In_a_Nutshell_Princeton_A._Zee_Quantum
PDF · 605 pages · 3.1 MB
Open PDF file
Princeton University Press graduate-level textbook by A. Zee, second edition, 2010. The extracted text shows front matter with reader praise and the full contents. Parts cover path integrals, Feynman diagrams, Dirac spinors, renormalization and gauge invariance, symmetry breaking, collective phenomena and condensed matter. It sits in Phil's downloaded physics books folder and is a published book, not Phil's own work.
AI-written summary; may contain errors.
Extracted text (machine-read; may contain errors)
Qluantum FieldTheory inaNutshell
‘Second Edition
p< ee
A ee]
Praise for the first edition
“Quantum field theory is an extraordinarily beautiful subject, but it can be an intimidating
one. The profound and deeply physical concepts it embodies can get lost, to the beginner,amidst its technicalities. In this book, Zee imparts the wisdom of an experienced andremarkably creative practitioner in a user-friendly style. I wish something like it had beenavailable when I was a student.”
—Frank Wilczek, Massachusetts Institute of Technology
“Finally! Zee has written a ground-breaking quantum field theory text based on the course
I made him teach when I chaired the Princeton physics department. With utmost clarityhe gives the eager student a light-hearted and easy-going introduction to the multifacetedwonders of quantum field theory. I wish I had this book when I taught the subject.”
—Marvin L. Goldberger, President, Emeritus, California Institute of Technology
“This book is filled with charming explanations that students will find beneficial.”
—Ed Witten, Institute for Advanced Study
“This book is perhaps the most user-friendly introductory text to the essentials of quantum
field theory and its many modern applications. With his physically intuitive approach,Professor Zee makes a serious topic more reachable for beginners, reducing the conceptualbarrier while preserving enough mathematical details necessary for a firm grasp of thesubject.”
—Bei Lok Hu, University of Maryland
“Like the famous Feynman Lectures on Physics, this book has the flavor of a good
blackboard lecture. Zee presents technical details, but only insofar as they serve the largerpurpose of giving insight into quantum field theory and bringing out its beauty.”
—Stephen M. Barr, University of Delaware
“This is a fantastic book—exciting, amusing, unique, and very valuable.”
—Clifford V . Johnson, University of Durham
“Tony Zee explains quantum field theory with a clear and engaging style. For budding or
seasoned condensed matter physicists alike, he shows us that field theory is a nourishing
nut to be cracked and savored.”
—Matthew P. A. Fisher, Kavli Institute for Theoretical Physics
“I was so engrossed that I spent all of Saturday and Sunday one weekend absorbing half
the book, to my wife’s dismay. Zee has a talent for explaining the most abstruse and
arcane concepts in an elegant way, using the minimum number of equations (the jokes
and anecdotes help)....I wish this were available when I was a graduate student. Buy
the book, keep it by your bed, and relish the insights delivered with such flair and grace.”
—N. P. Ong, Princeton University
What readers are saying
“Funny, chatty, physical: QFT education transformed!! This text stands apart from others
in so many ways that it’s difficult to list them all ....T h e exposition is breezy and chatty.
The text is never boring to read, and is at times very, very funny. Puns and jokes abound,as do anecdotes....A book which is much easier, and more fun, to read than any of the
others. Zee’s skills as a popular physics writer have been used to excellent effect in writingthis textbook....Wholeheartedly recommended.”
—M. Haque
“A readable, and rereadable instant classic on QFT ....A ta n introductory level, this type
of book—with its pedagogical (and often very funny) narrative—is priceless. [It] is fullof fantastic insights akin to reading the Feynman lectures. I have since used QFT in a
Nutshell as a review for [my] year-long course covering all of Peskin and Schroder, and
have been pleasantly surprised at how Zee is able to preemptively answer many of theopen questions that eluded me during my course ....I value QFT in a Nutshell the same
way I do the Feynman lectures....I t ’ sa text to teach an understanding of physics.”
—Flip Tanedo
“One of those books a person interested in theoretical physics simply must own! A real
scientific masterpiece. I bought it at the time I was a physics sophomore and that was thebest choice I could have made. It was this book that triggered my interest in quantum fieldtheory and crystallized my dreams of becoming a theoretical physicist....T h e main goal
of the book is to make the reader gain real intuition in the field. Amazin g...amusing...
real fun. What also distinguishes this book from others dealing with a similar subjectis that it is written like a tale....I feel enormously fortunate to have come across this
book at the beginning of my adventure with theoretical physics ....D e f initely the best
quantum field theory book I have ever read.”
—Anonymous
“I have used Quantum Field Theory in a Nutshell as the primary text....Ia m immensely
pleased with the book, and recommend it highly ....D o n ’ tl e tt h e‘ damn the torpedoes,
full steam ahead’ approach scare you off. Once you get used to seeing the physics quickly,I think you will find the experience very satisfying intellectually.”
—Jim Napolitano
“This is undoubtedly the best book I have ever read about the subject. Zee does a fantastic
job of explaining quantum field theory, in a way I have never seen before, and I have
read most of the other books on this topic. If you are looking for quantum field theoryexplanations that are clear, precise, concise, intuitive, and fun to read—this is the book
for you.”
—Anonymous
“One of the most artistic and deepest books ever written on quantum field theory.
Amazing...extremely pleasant...al o to f very deep and illuminating remarks....I
recommend the book by Zee to everybody who wants to get a clear idea what good physicsis about.”
—Slava Mukhanov
“Perfect for learning field theory on your own—by far the clearest and easiest to follow
book I’ve found on the subject.”
—Ian Z. Lovejoy
“A beautifully written introduction to the modern view of field s...breezy and
enchanting, leading to exceptional clarity without sacrificing depth, breadth, or rigorof content....[It] passes my test of true greatness: I wish it had been the first book onthis topic that I had found.”
—Jeffrey D. Scargle
“A breeze of fresh air...a real literary gem which will be useful for students who make
their first steps in this difficult subject and an enjoyable treat for experts, who will findnew and deep insights. Indeed, the Nutshell is like a bright light source shining among
tall and heavy trees—the many more formal books that exist—and helps seeing the forestas a whole!...I have been practicing QFT during the past two decades and with all my
experience I was thrilled with enjoyment when I read some of the sections.”
—Joshua Feinberg
“This text not only teaches up-to-date quantum field theory, but also tells readers how
research is actually done and shows them how to think about physics. [It teaches thingsthat] people usually say ‘cannot be learned from books.’ [It is] in the same style as Fearful
Symmetry and Einstein ’s Universe . All three books...a r e classics.”
—Yu Shi
“I belong to the [group of ] enthusiastic laymen having enough curiosity and insistence...
but lacking the mastery of advanced math and physics....I really could not see the forest
for the trees. But at long last I got this book!”
—Makay Attila
“More fun than any other QFT book I have read. The comparisons to Feynman’s
writings made by several of the reviewers seem quite apt....H i s enthusiasm is quite
infectious....I doubt that any other book will spark your interest like this one does.”
—Stephen Wandzura
“I’m having a blast reading this book. It’s both deep and entertaining; this is a rare breed,
indeed. I usually prefer the more formal style (big Landau fan), but I have to say that whenZee has the talent to present things his way, it’s a definite plus.”
—Pierre Jouvelot
“Required reading for QFT: [it] heralds the introduction of a book on quantum field theory
that you can sit down and read. My professor’s lectures made much more sense as Ifollowed along in this book, because concepts were actually EXPLAINED, not just workedout.”
—Alexander Scott
“Not your father’s quantum field theory text: I particularly appreciate that things are
motivated physically before their mathematical articulation....M o s t especially though,
the author’s ‘heuristic’ descriptions are the best I have read anywhere. From them alonethe essential ideas become crystal clear.”
—Dan Dill
Q uantum Field Theory in a Nutshell
This page intentionally left blank
Quantum Field Theory in a Nutshell
SECOND EDITION
A. Zee
PRINCETON UNIVERSITY PRESS.PRINCETON AND OXFORD
Copyright © 2010 by Princeton University Press
Published by Princeton University Press, 41 William Street,
Princeton, New Jersey 08540In the United Kingdom: Princeton University Press,6 Oxford Street, Woodstock, Oxfordshire OX20 1TW
All Rights ReservedLibrary of Congress Cataloging-in-Publication Data
Zee, A.
Quantum field theory in a nutshell / A. Zee.—2nd ed.
p. cm.
Includes bibliographical references and index.ISBN 978-0-691-14034-6 (hardcover : alk. paper) 1. Quantum field theory. I. Title.QC174.45.Z44 2010530.14
/prime3—dc22 2009015469
British Library Cataloging-in-Publication Data is availableThis book has been composed in Scala LF with ZzT EX
by Princeton Editorial Associates, Inc., Scottsdale, Arizona
Printed on acid-free paper.press.princeton.eduPrinted in the United States of America1 0987654321
T o my parents,
who valued education above all else
This page intentionally left blank
Contents
Preface to the First Edition xv
Preface to the Second Edition xix
Convention, Notation, and Units xxv
IPart I: Motivation and Foundation
I.1 Who Needs It? 3
I.2 Path Integral Formulation of Quantum Physics 7
I.3 From Mattress to Field 17
I.4 From Field to Particle to Force 26
I.5 Coulomb and Newton: Repulsion and Attraction 32
I.6 Inverse Square Law and the Floating 3-Brane 40
I.7 Feynman Diagrams 43
I.8 Quantizing Canonically 61
I.9 Disturbing the Vacuum 70
I.10 Symmetry 76
I.11 Field Theory in Curved Spacetime 81
I.12 Field Theory Redux 88
IIPart II: Dirac and the Spinor
II.1 The Dirac Equation 93
II.2 Quantizing the Dirac Field 107
II.3 Lorentz Group and Weyl Spinors 114
II.4 Spin-Statistics Connection 120
xii | Contents
II.5 Vacuum Energy, Grassmann Integrals, and Feynman Diagrams
for Fermions 123
II.6 Electron Scattering and Gauge Invariance 132
II.7 Diagrammatic Proof of Gauge Invariance 144
II.8 Photon-Electron Scattering and Crossing 152
III Part III: Renormalization and Gauge Invariance
III.1 Cutting Off Our Ignorance 161
III.2 Renormalizable versus Nonrenormalizable 169
III.3 Counterterms and Physical Perturbation Theory 173
III.4 Gauge Invariance: A Photon Can Find No Rest 182
III.5 Field Theory without Relativity 190
III.6 The Magnetic Moment of the Electron 194
III.7 Polarizing the Vacuum and Renormalizing the Charge 200
III.8 Becoming Imaginary and Conserving Probability 207
IV Part IV: Symmetry and Symmetry Breaking
IV .1 Symmetry Breaking 223
IV .2 The Pion as a Nambu-Goldstone Boson 231
IV .3 Effective Potential 237
IV .4 Magnetic Monopole 245
IV .5 Nonabelian Gauge Theory 253
IV .6 The Anderson-Higgs Mechanism 263
IV .7 Chiral Anomaly 270
VPart V: Field Theory and Collective Phenomena
V .1 Superfluids 283
V .2 Euclid, Boltzmann, Hawking, and Field Theory at Finite Temperature 287
V .3 Landau-Ginzburg Theory of Critical Phenomena 292
V .4 Superconductivity 295
V .5 Peierls Instability 298
V .6 Solitons 302
V .7 Vortices, Monopoles, and Instantons 306
VI Part VI: Field Theory and Condensed Matter
VI.1 Fractional Statistics, Chern-Simons Term, and Topological
Field Theory 315
VI.2 Quantum Hall Fluids 322
Contents | xiii
VI.3 Duality 331
VI.4 The σModels as Effective Field Theories 340
VI.5 Ferromagnets and Antiferromagnets 344
VI.6 Surface Growth and Field Theory 347
VI.7 Disorder: Replicas and Grassmannian Symmetry 350
VI.8 Renormalization Group Flow as a Natural Concept in High Energy
and Condensed Matter Physics 356
VII Part VII: Grand Unification
VII.1 Quantizing Yang-Mills Theory and Lattice Gauge Theory 371
VII.2 Electroweak Unification 379
VII.3 Quantum Chromodynamics 385
VII.4 Large NExpansion 394
VII.5 Grand Unification 407
VII.6 Protons Are Not Forever 413
VII.7 SO(10) Unification 421
VIII Part VIII: Gravity and Beyond
VIII.1 Gravity as a Field Theory and the Kaluza-Klein Picture 433
VIII.2 The Cosmological Constant Problem and the Cosmic Coincidence
Problems 448
VIII.3 Effective Field Theory Approach to Understanding Nature 452
VIII.4 Supersymmetry: A Very Brief Introduction 461
VIII.5 A Glimpse of String Theory as a 2-Dimensional Field Theory 469
Closing Words 473
NPart N
N.1 Gravitational Waves and Effective Field Theory 479
N.2 Gluon Scattering in Pure Yang-Mills Theory 483
N.3 Subterranean Connections in Gauge Theories 497
N.4 Is Einstein Gravity Secretly the Square of Yang-Mills Theory? 513
More Closing Words 521
Appendix A: Gaussian Integration and the Central Identity of Quantum
Field Theory 523
Appendix B: A Brief Review of Group Theory 525
xiv | Contents
Appendix C: Feynman Rules 534
Appendix D: Various Identities and Feynman Integrals 538
Appendix E: Dotted and Undotted Indices and the Majorana Spinor 541
Solutions to Selected Exercises 545
Further Reading 559
Index 563
Preface to the First Edition
As a student, I was rearing at the bit, after a course on quantum mechanics, to learn
quantum field theory, but the books on the subject all seemed so formidable. Fortunately,I came across a little book by Mandl on field theory, which gave me a taste of the subjectenabling me to go on and tackle the more substantive texts. I have since learned that otherphysicists of my generation had similar good experiences with Mandl.
In the last three decades or so, quantum field theory has veritably exploded and Mandl
would be hopelessly out of date to recommend to a student now. Thus I thought of writinga book on the essentials of modern quantum field theory addressed to the bright and eagerstudent who has just completed a course on quantum mechanics and who is impatient tostart tackling quantum field theory.
I envisaged a relatively thin book, thin at least in comparison with the many weighty
tomes on the subject. I envisaged the style to be breezy and colloquial, and the choiceof topics to be idiosyncratic, certainly not encyclopedic. I envisaged having many shortchapters, keeping each chapter “bite-sized.”
The challenge in writing this book is to keep it thin and accessible while at the same
time introducing as many modern topics as possible. A tough balancing act! In the end,I had to be unrepentantly idiosyncratic in what I chose to cover. Note to the prospectivebook reviewer: You can always criticize the book for leaving out your favorite topics. I donot apologize in any way, shape, or form. My motto in this regard (and in life as well),taken from the Ricky Nelson song “Garden Party,” is “You can’t please everyone so yougotta please yourself.”
This book differs from other quantum field theory books that have come out in recent
years in several respects.
I want to get across the important point that the usefulness of quantum field theory is far
from limited to high energy physics, a misleading impression my generation of theoreticalphysicists were inculcated with and which amazingly enough some recent textbooks on
xvi | Preface to the First Edition
quantum field theory (all written by high energy physicists) continue to foster. For instance,
the study of driven surface growth provides a particularly clear, transparent, and physicalexample of the importance of the renormalization group in quantum field theory. Insteadof being entangled in all sorts of conceptual irrelevancies such as divergences, we havethe obviously physical notion of changing the ruler used to measure the fluctuatingsurface. Other examples include random matrix theory and Chern-Simons gauge theoryin quantum Hall fluids. I hope that condensed matter theory students will find this bookhelpful in getting a first taste of quantum field theory. The book is divided into eight parts,
1
with two devoted more or less exclusively to condensed matter physics.
I try to give the reader at least a brief glimpse into contemporary developments, for
example, just enough of a taste of string theory to whet the appetite. This book is perhapsalso exceptional in incorporating gravity from the beginning. Some topics are treated quitedifferently than in traditional texts. I introduce the Faddeev-Popov method to quantizeelectromagnetism and the language of differential forms to develop Yang-Mills theory, forexample.
The emphasis is resoundingly on the conceptual rather than the computational. The
only calculation I carry out in all its gory details is that of the magnetic moment of theelectron. Throughout, specific examples rather than heavy abstract formalism will befavored. Instead of dealing with the most general case, I always opt for the simplest.
I had to struggle constantly between clarity and wordiness. In trying to anticipate and to
minimize what would confuse the reader, I often find that I have to belabor certain pointsmore than what I would like.
I tried to avoid the dreaded phrase “It can be shown tha t...”a s much as possible.
Otherwise, I could have written a much thinner book than this! There are indeed thinnerbooks on quantum field theory: I looked at a couple and discovered that they hardly explainanything. I must confess that I have an almost insatiable desire to explain.
As the manuscript grew, the list of topics that I reluctantly had to drop also kept growing.
So many beautiful results, but so little space! It almost makes me ill to think about all thestuff (bosonization, instanton, conformal field theory, etc., etc.) I had to leave out. As onecolleague remarked, the nutshell is turning into a coconut shell!
Shelley Glashow once described the genesis of physical theories: “Tapestries are made
by many artisans working together. The contributions of separate workers cannot bediscerned in the completed work, and the loose and false threads have been covered over.” Iregret that other than giving a few tidbits here and there I could not go into the fascinatinghistory of quantum field theory, with all its defeats and triumphs. On those occasionswhen I refer to original papers I suffer from that disconcerting quirk of human psychologyof tending to favor my own more than decorum might have allowed. I certainly did notattempt a true bibliography.
1Murray Gell-Mann used to talk about the eightfold way to wisdom and salvation in Buddhism (M. Gell-Mann
and Y . Ne’eman, The Eightfold Way ). Readers familiar with contemporary Chinese literature would know that the
celestial dragon has eight parts.
Preface to the First Edition | xvii
The genesis of this book goes back to the quantum field theory course I taught as a
beginning assistant professor at Princeton University. I had the enormous good fortuneof having Ed Witten as my teaching assistant and grader. Ed produced lucidly writtensolutions to the homework problems I assigned, to the extent that the next year I wentto the chairman to ask “What is wrong with the TA I have this year? He is not half asgood as the guy last year!” Some colleagues asked me to write up my notes for a muchneeded text (those were the exciting times when gauge theories, asymptotic freedom,and scores of topics not to be found in any texts all had to be learned somehow) but awiser senior colleague convinced me that it might spell disaster for my research career.Decades later, the time has come. I particularly thank Murph Goldberger for urging meto turn what expository talents I have from writing popular books to writing textbooks. Itis also a pleasure to say a word in memory of the late Sam Treiman, teacher, colleague,and collaborator, who as a member of the editorial board of Princeton University Presspersuaded me to commit to this project. I regret that my slow pace in finishing the bookdeprived him of seeing the finished product.
Over the years I have refined my knowledge of quantum field theory in discussions
with numerous colleagues and collaborators. As a student, I attended courses on quan-tum field theory offered by Arthur Wightman, Julian Schwinger, and Sidney Coleman. Iwas fortunate that these three eminent physicists each has his own distinctive style andapproach.
The book has been tested “in the field” in courses I taught. I used it in my field theory
course at the University of California at Santa Barbara, and I am grateful to some ofthe students, in particular Ted Erler, Andrew Frey, Sean Roy, and Dean Townsley, forcomments. I benefitted from the comments of various distinguished physicists who readall or parts of the manuscript, including Steve Barr, Doug Eardley, Matt Fisher, MurphGoldberger, Victor Gurarie, Steve Hsu, Bei-lok Hu, Clifford Johnson, Mehran Kardar, IanLow, Joe Polchinski, Arkady Vainshtein, Frank Wilczek, Ed Witten, and especially JoshuaFeinberg. Joshua also did many of the exercises.
Talking about exercises: You didn’t get this far in physics without realizing the absolute
importance of doing exercises in learning a subject. It is especially important that you domost of the exercises in this book, because to compensate for its relative slimness I haveto develop in the exercises a number of important points some of which I need for laterchapters. Solutions to some selected problems are given.
I will maintain a web page http://theory.kitp.ucsb.edu/~zee/nuts.html listing all the
errors, typographical and otherwise, and points of confusion that will undoubtedly cometo my attention.
I thank my editors, Trevor Lipscombe, Sarah Green, and the staff of Princeton Editorial
Associates (particularly Cyd Westmoreland and Evelyn Grossberg) for their advice and forseeing this project through. Finally, I thank Peter Zee for suggesting the cover painting.
This page intentionally left blank
Preface to the Second Edition
What one fool could understand, another can.
—R. P. Feynman1
Appreciating the appreciators
It has been nearly six years since this book was published on March 10, 2003. Since authorsoften think of books as their children, I may liken the flood of appreciation from readers,students, and physicists to the glorious report cards a bright child brings home fromschool. Knowing that there are people who appreciate the care and clarity crafted into thepedagogy is a most gratifying feeling. In working on this new edition, merely looking atthe titles of the customer reviews on Amazon.com would lighten my task and quicken mypace: “Funny, chatty, physical. QFT education transformed!,” “A readable, and re-readableinstant classic on QFT ,” “A must read book if you want to understand essentials in QFT ,”“One of the most artistic and deepest books ever written on quantum field theory,” “Perfectfor learning field theory on your own,” “Both deep and entertaining,” “One of those booksa person interested in theoretical physics simply must own,” and so on.
In a Physics T oday review, Zvi Bern, a preeminent younger field theorist, wrote:
Perhaps foremost in his mind was how to make Quantum Field Theory in a Nutshell as much fun
as possible....I have not had this much fun with a physics book since reading The Feynman
Lectures on Physics ....[This is a book] that no student of quantum field theory should be
without. Quantum Field Theory in a Nutshell is the ideal book for a graduate student to curl up
with after having completed a course on quantum mechanics. But, mainly, it is for anyone whowishes to experience the sheer beauty and elegance of quantum field theory.
A classical Chinese scholar famously lamented “He who knows me are so few!” but here
Zvi read my mind.
Einstein proclaimed, “Physics should be made as simple as possible, but not any
simpler.” My response would be “Physics should be made as fun as possible, but not
1R. P. Feynman, QED: The Strange Theory of Light and Matter, p. xx.
xx | Preface to the Second Edition
any funnier.” I overcame the editor’s reluctance and included jokes and stories. And yes, I
have also written a popular book Fearful Symmetry about the “sheer beauty and elegance”
of modern physics, which at least in that book largely meant quantum field theory. I wantto share that sense of fun and beauty as much as possible. I’ve heard some people say that“Beauty is truth” but “Beauty is fun” is more like it.
I had written books before, but this was my first textbook. The challenges and rewards
in writing different types of book are certainly different, but to me, a university professordevoted to the ideals of teaching, the feeling of passing on what I have learned andunderstood is simply incomparable. (And the nice part is that I don’t have to hand outfinal grades.) It may sound corny, but I owe it, to those who taught me and to thoseauthors whose field theory texts I studied, to give something back to the theoretical physicscommunity. It is a wonderful feeling for me to meet young hotshot researchers who hadstudied this text and now know more about field theory than I do.
How I made the book better: The first text that covers the twenty-first century
When my editor Ingrid Gnerlich asked me for a second edition I thought long and hardabout how to make this edition better than the first. I have clarified and elaborated hereand there, added explanations and exercises, and done more “practical” Feynman diagramcalculations to appease those readers of the first edition who felt that I didn’t calculateenough. There are now three more chapters in the main text. I have also made the “mostaccessible” text on quantum field theory even more accessible by explaining stuff thatI thought readers who already studied quantum mechanics should know. For example,I added a concise review of the Dirac delta function to chapter I.2. But to the guy onAmazon.com who wanted complex analysis explained, sorry, I won’t do it. There is a limit.Already, I gave a basically self-contained coverage of group theory.
More excitingly, and to make my life more difficult, I added, to the existing eight parts
(of the celestial dragon), a new part consisting of four chapters, covering field theoretichappenings of the last decade or so. Thus I can say that this is the first text since the birthof quantum field theory in the late 1920s that covers the twenty-first century.
Quantum field theory is a mature but certainly not a finished subject, as some stu-
dents mistakenly believe. As one of the deepest constructs in theoretical physics and allencompassing in its reach, it is bound to have yet unplumbed depths, secret subterraneanconnections, and delightful surprises. While many theoretical physicists have moved pastquantum field theory to string theory and even string field theory, they often take the limitin which the string description reduces to a field description, thus on occasion revealingpreviously unsuspected properties of quantum field theories. We will see an example inchapter N.4.
My friends admonished me to maintain, above all else, the “delightful tone” of the first
edition. I hope that I have succeeded, even though the material contained in part N is “hotoff the stove” stuff, unlike the long-understood material covered in the main text. I alsoadded a few jokes and stories, such as the one about Fermi declining to trace.
Preface to the Second Edition | xxi
As with the first edition, I will maintain a web site http://theory.kitp.ucsb.edu/~zee/
nuts2.html listing the errors, typographical or otherwise, that will undoubtedly come tomy attention.
Encouraging words
In the quote that started this preface, Feynman was referring to himself, and to you! Ofcourse, Feynman didn’t simply understand the quantum field theory of electromagnetism,he also invented a large chunk of it. To paraphrase Feynman, I wrote this book for foolslike you and me. If a fool like me could write a book on quantum field theory, then surelyyou can understand it.
As I said in the preface to the first edition, I wrote this book for those who, having
learned quantum mechanics, are eager to tackle quantum field theory. During a sabbaticalyear (2006–07) I spent at Harvard, I was able to experimentally verify my hypothesis thata person who has mastered quantum mechanics could understand my book on his or herown without much difficulty. I was sent a freshman who had taught himself quantummechanics in high school. I gave him my book to read and every couple of weeks or sohe came by to ask a question or two. Even without these brief sessions, he would haveunderstood much of the book. In fact, at least half of his questions stem from the holesin his knowledge of quantum mechanics. I have incorporated my answers to his fieldtheoretic questions into this edition.
As I also said in the original preface, I had tested some of the material in the book “in the
field” in courses I taught at Princeton University and later at the University of California atSanta Barbara. Since 2003, I have been gratified to know that it has been used successfullyin courses at many institutions.
I understand that, of the different groups of readers, those who are trying to learn
quantum field theory on their own could easily get discouraged. Let me offer you somecheering words. First of all, that is very admirable of you! Of all the established subjectsin theoretical physics, quantum field theory is by far the most subtle and profound. Byconsensus it is much much harder to learn than Einstein’s theory of gravity, which in factshould properly be regarded as part of field theory, as will be made clear in this book. Sodon’t expect easy cruising, particularly if you don’t have someone to clarify things for youonce in a while. Try an online physics forum. Do at least some of the exercises. Remember:“No one expects a guitarist to learn to play by going to concerts in Central Park or byspending hours reading transcriptions of Jimi Hendrix solos. Guitarists practice. Guitaristsplay the guitar until their fingertips are calloused. Similarly, physicists solve problems.”
2
Of course, if you don’t have the prerequisites, you won’t be able to understand this or anyother field theory text. But if you have mastered quantum mechanics, keep on truckingand you will get there.
2N. Newbury et al., Princeton Problems in Physics with Solutions, Princeton University Press, Princeton, 1991.
xxii | Preface to the Second Edition
The view will be worth it, I promise. My thesis advisor Sidney Coleman used to start his
field theory course saying, “Not only God knows, I know, and by the end of the semester,you will know.” By the end of this book, you too will know how God weaves the universeout of a web of interlocking fields. I would like to change Dirac’s statement “God is amathematician” to “God is a quantum field theorist.”
Some of you steady truckers might want to ask what to do when you get to the end. Dur-
ing my junior year in college, after my encounter with Mandl, I asked Arthur Wightmanwhat to read next. He told me to read the textbook by S. S. Schweber, which at close to athousand pages was referred to by students as “the monster” and which could be extremelyopaque at places. After I slugged my way to the end, Wightman told me, “Read it again.”Fortunately for me, volume I of Bjorken and Drell had already come out. But there is wis-dom in reading a book again; things that you miss the first time may later leap out at you.So my advice is “Read it again.” Of course, every physics student also knows that differentexplanations offered by different books may click with different folks. So read other fieldtheory books. Quantum field theory is so profound that most people won’t get it in onepass.
On the subject of other field theory texts: James Bjorken kindly wrote in my much-used
copy of Bjorken and Drell that the book was obsolete. Hey BJ, it isn’t. Certainly, volume Iwill never be pass ´e. On another occasion, Steve Weinberg told me, referring to his field
theory book, that “I wrote the book that I would have liked to learn from.” I could equallywell say that “I wrote the book that Iwould have liked to learn from.” Without the least
bit of hubris, I can say that I prefer my book to Schweber’s. The moral here is that if youdon’t like this book you should write your own.
I try not to do clunky
I explained my philosophy in the preface to the first edition, but allow me a few morewords here. I will teach you how to calculate, but I also have what I regard as a higher aim,to convey to you an enjoyment of quantum field theory in all its splendors (and by “all” Imean not merely quantum field theory as defined by some myopic physicists as applicableonly to particle physics). I try to erect an elegant and logically tight framework and put alight touch on a heavy subject.
In spite of the image conjured up by Zvi Bern of some future field theorist curled up
in bed reading this book, I expect you to grab pen and paper and work. You could doit in bed if you want, but work you must. I intentionally did not fill in all the steps; itwould hardly be a light touch if I do every bit of algebra for you. Nevertheless, I have donealgebra when I think that it would help you. Actually, I love doing algebra, particularlywhen things work out so elegantly as in quantum field theory. But I don’t do clunky. Ido not like clunky-looking equations. I avoid spelling everything out and so expect you
to have a certain amount of “sense.” As a small example, near the end of chapter I.10 Isuppressed the spacetime dependence of the fields ϕ
aandδϕa. If you didn’t realize, after
Preface to the Second Edition | xxiii
some 70 pages, that fields are functions of where you are in spacetime, you are quite lost,
my friend. My plan is to “keep you on your toes” and I purposely want you to feel puzzledoccasionally. I have faith that the sort of person who would be reading this book can alwaysfigure it out after a bit of thought. I realize that there are at least three distinct groups ofreaders, but let me say to the students, “How do you expect to do research if you have tobe spoon-fed from line to line in a textbook?”
Nuts who do not appreciate the Nutshell
In the original preface, I quoted Ricky Nelson on the impossibility of pleasing everyone and
so I was not at all surprised to find on Amazon.com a few people whom one of my friendscalls “nuts who do not appreciate the Nutshell .” My friends advise me to leave these people
alone but I am sufficiently peeved to want to say a few words in my defense, no matter hownutty the charge. First, I suppose that those who say the book is too mathematical cancelout those who say the book is not mathematical enough. The people in the first group arenot informed, while those in the second group are misinformed.
Quantum field theory does not have to be mathematical. I know of at least three Field
Medalists who enjoyed the book. A review for the American Mathematical Society offeredthis deep statement in praise of the book: “It is often deeper to know why something istrue rather than to have a proof that it is true.” (Indeed, a Fields Medalist once told me thattop mathematicians secretly think like physicists and after they work out the broad outlineof a proof they then dress it up with epsilons and deltas. I have no idea if this is true onlyfor one, for many, or for all Fields Medalists. I suspect that it is true for many.)
Then there is the person who denounces the book for its lack of rigor. Well, I happen to
know, or at least used to know, a thing or two about mathematical rigor, since I wrote mysenior thesis with Wightman on what I would call “fairly rigorous” quantum field theory.As we like to say in the theoretical physics community, too much rigor soon leads to rigormortis. Be warned. Indeed, as Feynman would tell students, if this ain’t rigorous enoughfor you the math department is just one building over. So read a more rigorous book. It isa free country.
More serious is the impression that several posters on Amazon.com have that the book is
too elementary. I humbly beg to differ. The book gives the impression of being elementarybut in fact covers more material than many other texts. If you master everything in theNutshell, you would know more than most professors of field theory and could start doing
research. I am not merely making an idle claim but could give an actual proof. All theingredients that went into the spinor helicity formalism that led to a deep field theoreticdiscovery described in part N could be found in the first edition of this book. Of course,reading a textbook is not enough; you have to come up with the good ideas.
As for he who says that the book does not look complicated enough and hence can’t be
a serious treatment, I would ask him to compare a modern text on electromagnetism withMaxwell’s treatises.
xxiv | Preface to the Second Edition
Thanks
In the original preface and closing words, I mentioned that I learned a great deal of quan-
tum field theory from Sidney Coleman. His clarity of thought and lucid exposition havealways inspired me. Unhappily, he passed away in 2007. After this book was published, Ivisited Sidney on different occasions, but sadly, he was already in a mental fog.
In preparing this second edition, I am grateful to Nima Arkani-Hamed, Yoni Ben-Tov,
Nathan Berkovits, Marty Einhorn, Joshua Feinberg, Howard Georgi, Tim Hsieh, BrendanKeller, Joe Polchinski, Yong-shi Wu, and Jean-Bernard Zuber for their helpful comments.Some of them read parts or all of the added chapters. I thank especially Zvi Bern andRafael Porto for going over the chapters in part N with great care and for many usefulsuggestions. I also thank Craig Kunimoto, Richard Neher, Matt Pillsbury, and RafaelPorto for teaching me the black art of composing equations on the computer. My editorat Princeton University Press, Ingrid Gnerlich, has always been a pleasure to talk toand work with. I also thank Kathleen Cioffi and Cyd Westmoreland for their meticulouswork in producing this book. Last but not least, I am grateful to my wife Janice for herencouragement and loving support.
Convention, Notation, and Units
For the same reason that we no longer use a certain king’s feet to measure distance, we use
natural units in which the speed of light cand the Dirac symbol /planckover2piare both set equal to 1.
Planck made the profound observation that in natural units all physical quantities can beexpressed in terms of the Planck mass M
Planck≡1//radicalbig
GNewton /similarequal1019Gev. The quantities
cand/planckover2piare not so much fundamental constants as conversion factors. In this light, I am
genuinely puzzled by condensed matter physicists carrying around Boltzmann’s constantk, which is no different from the conversion factor between feet and meters.
Spacetime coordinates x
μare labeled by Greek indices (μ =0, 1, 2, 3 ) with the time
coordinate x0sometimes denoted by t. Space coordinates xiare labeled by Latin indices
(i=1, 2, 3 ) and ∂μ≡∂/∂xμ. We use a Minkowski metric ημνwith signature ( +,−,−,−)
so that η00=+ 1. We write ημν∂μϕ∂νϕ=∂μϕ∂μϕ=(∂ϕ)2=(∂ϕ/∂t)2−/summationtext
i(∂ϕ/∂xi)2. The
metric in curved spacetime is always denoted by gμν, but often I will also use gμνfor the
Minkowski metric when the context indicates clearly that we are in flat spacetime.
Since I will be talking mostly about relativistic quantum field theory in this book I will
without further clarification use a relativistic language. Thus, when I speak of momentum,unless otherwise specified, I mean energy and momentum. Also since /planckover2pi=1, I will not
distinguish between wave vector kand momentum, and between frequency ωand energy.
In local field theory I deal primarily with the Lagrangian density Land not the La-
grangian L=/integraltext
d
3xL. As is common practice in the literature and in oral discussion, I
will often abuse terminology and simply refer to Las the Lagrangian. I will commit other
minor abuses such as writing 1 instead of Ifor the unit matrix. I use the same symbol
ϕfor the Fourier transform ϕ(k) of a function ϕ(x) whenever there is no risk of confu-
sion, as is almost always the case. I prefer an abused terminology to cluttered notation and
unbearable pedantry.
The symbol ∗denotes complex conjugation, and † hermitean conjugation: The former
applies to a number and the latter to an operator. I also use the notation c.c. and h.c. Often
xxvi | Convention, Notation, and Units
when there is no risk of confusion I abuse the notation, using † when I should use ∗.F o r
instance, in a path integral, bosonic fields are just number-valued fields, but neverthelessI write ϕ
†rather than ϕ∗. For a matrix M, then of course M†andM∗should be carefully
distinguished from each other.
I made an effort to get factors of 2 and πright, but some errors will be inevitable.
Q uantum Field Theory in a Nutshell
This page intentionally left blank
Part I Motivation and Foundation
This page intentionally left blank
I.1 Who Needs It?
Who needs quantum field theory?
Quantum field theory arose out of our need to describe the ephemeral nature of life.
No, seriously, quantum field theory is needed when we confront simultaneously the two
great physics innovations of the last century of the previous millennium: special relativityand quantum mechanics. Consider a fast moving rocket ship close to light speed. You needspecial relativity but not quantum mechanics to study its motion. On the other hand, tostudy a slow moving electron scattering on a proton, you must invoke quantum mechanics,but you don’t have to know a thing about special relativity.
It is in the peculiar confluence of special relativity and quantum mechanics that a new
set of phenomena arises: Particles can be born and particles can die. It is this matter ofbirth, life, and death that requires the development of a new subject in physics, that ofquantum field theory.
Let me give a heuristic discussion. In quantum mechanics the uncertainty principle tells
us that the energy can fluctuate wildly over a small interval of time. According to specialrelativity, energy can be converted into mass and vice versa. With quantum mechanics andspecial relativity, the wildly fluctuating energy can metamorphose into mass, that is, intonew particles not previously present.
Write down the Schr ¨odinger equation for an electron scattering off a proton. The
equation describes the wave function of one electron, and no matter how you shakeand bake the mathematics of the partial differential equation, the electron you followwill remain one electron. But special relativity tells us that energy can be converted tomatter: If the electron is energetic enough, an electron and a positron (“the antielectron”)can be produced. The Schr ¨odinger equation is simply incapable of describing such a
phenomenon. Nonrelativistic quantum mechanics must break down.
You saw the need for quantum field theory at another point in your education. T oward
the end of a good course on nonrelativistic quantum mechanics the interaction betweenradiation and atoms is often discussed. You would recall that the electromagnetic field is
4 | I. Motivation and Foundation
Figure I.1.1
treated as a field; well, it is a field. Its Fourier components are quantized as a collection
of harmonic oscillators, leading to creation and annihilation operators for photons. Sothere, the electromagnetic field is a quantum field. Meanwhile, the electron is treated as apoor cousin, with a wave function /Psi1(x) governed by the good old Schr ¨odinger equation.
Photons can be created or annihilated, but not electrons. Quite aside from the experimentalfact that electrons and positrons could be created in pairs, it would be intellectually moresatisfying to treat electrons and photons, as they are both elementary particles, on the samefooting.
So, I was more or less right: Quantum field theory is a response to the ephemeral nature
of life.
All of this is rather vague, and one of the purposes of this book is to make these remarks
more precise. For the moment, to make these thoughts somewhat more concrete, let usask where in classical physics we might have encountered something vaguely resemblingthe birth and death of particles. Think of a mattress, which we idealize as a 2-dimensionallattice of point masses connected to each other by springs (fig. I.1.1). For simplicity, letus focus on the vertical displacement [which we denote by q
a(t)] of the point masses and
neglect the small horizontal movement. The index asimply tells us which mass we are
talking about. The Lagrangian is then
L=1
2(/summationdisplay
am˙q2
a−/summationdisplay
a,bkabqaqb−/summationdisplay
a,b,cgabcqaqbqc−...) (1)
Keeping only the terms quadratic in q(the “harmonic approximation”) we have the equa-
tions of motion m¨qa=−/summationtext
bkabqb. T aking the q’s as oscillating with frequency ω,w e
have/summationtext
bkabqb=mω2qa. The eigenfrequencies and eigenmodes are determined, respec-
tively, by the eigenvalues and eigenvectors of the matrix k. As usual, we can form wave
packets by superposing eigenmodes. When we quantize the theory, these wave packets be-have like particles, in the same way that electromagnetic wave packets when quantizedbehave like particles called photons.
I.1. Who Needs It? | 5
Since the theory is linear, two wave packets pass right through each other. But once we
include the nonlinear terms, namely the terms cubic, quartic, and so forth in the q’s in
(1), the theory becomes anharmonic. Eigenmodes now couple to each other. A wave packetmight decay into two wave packets. When two wave packets come near each other, theyscatter and perhaps produce more wave packets. This naturally suggests that the physicsof particles can be described in these terms.
Quantum field theory grew out of essentially these sorts of physical ideas.It struck me as limiting that even after some 75 years, the whole subject of quantum
field theory remains rooted in this harmonic paradigm, to use a dreadfully pretentiousword. We have not been able to get away from the basic notions of oscillations and wavepackets. Indeed, string theory, the heir to quantum field theory, is still firmly founded onthis harmonic paradigm. Surely, a brilliant young physicist, perhaps a reader of this book,will take us beyond.
Condensed matter physics
In this book I will focus mainly on relativistic field theory, but let me mention here thatone of the great advances in theoretical physics in the last 30 years or so is the increasinglysophisticated use of quantum field theory in condensed matter physics. At first sight thisseems rather surprising. After all, a piece of “condensed matter” consists of an enormousswarm of electrons moving nonrelativistically, knocking about among various atomic ionsand interacting via the electromagnetic force. Why can’t we simply write down a giganticwave function /Psi1(x
1,x2,... ,xN), where xjdenotes the position of the jth electron and N
is a large but finite number? Okay, /Psi1is a function of many variables but it is still governed
by a nonrelativistic Schr ¨odinger equation.
The answer is yes, we can, and indeed that was how solid state physics was first studied
in its heroic early days (and still is in many of its subbranches).
Why then does a condensed matter theorist need quantum field theory? Again, let us
first go for a heuristic discussion, giving an overall impression rather than all the details. Ina typical solid, the ions vibrate around their equilibrium lattice positions. This vibrationaldynamics is best described by so-called phonons, which correspond more or less to thewave packets in the mattress model described above.
This much you can read about in any standard text on solid state physics. Furthermore,
if you have had a course on solid state physics, you would recall that the energy levelsavailable to electrons form bands. When an electron is kicked (by a phonon field say) froma filled band to an empty band, a hole is left behind in the previously filled band. Thishole can move about with its own identity as a particle, enjoying a perfectly comfortableexistence until another electron comes into the band and annihilates it. Indeed, it was
with a picture of this kind that Dirac first conceived of a hole in the “electron sea” as theantiparticle of the electron, the positron.
We will flesh out this heuristic discussion in subsequent chapters in parts V and VI.
6 | I. Motivation and Foundation
Marriages
T o summarize, quantum field theory was born of the necessity of dealing with the marriage
of special relativity and quantum mechanics, just as the new science of string theory isbeing born of the necessity of dealing with the marriage of general relativity and quantummechanics.
I.2 Path Integral Formulation of Quantum Physics
The professor’s nightmare: a wise guy in the class
As I noted in the preface, I know perfectly well that you are eager to dive into quantum field
theory, but first we have to review the path integral formalism of quantum mechanics. Thisformalism is not universally taught in introductory courses on quantum mechanics, buteven if you have been exposed to it, this chapter will serve as a useful review. The reason Istart with the path integral formalism is that it offers a particularly convenient way of goingfrom quantum mechanics to quantum field theory. I will first give a heuristic discussion,to be followed by a more formal mathematical treatment.
Perhaps the best way to introduce the path integral formalism is by telling a story,
certainly apocryphal as many physics stories are. Long ago, in a quantum mechanics class,the professor droned on and on about the double-slit experiment, giving the standardtreatment. A particle emitted from a source S(fig. I.2.1) at time t=0 passes through one
or the other of two holes, A
1andA2, drilled in a screen and is detected at time t=Tby
a detector located at O. The amplitude for detection is given by a fundamental postulate
of quantum mechanics, the superposition principle, as the sum of the amplitude for theparticle to propagate from the source Sthrough the hole A
1and then onward to the point
Oand the amplitude for the particle to propagate from the source Sthrough the hole A2
and then onward to the point O.
Suddenly, a very bright student, let us call him Feynman, asked, “Professor, what if
we drill a third hole in the screen?” The professor replied, “Clearly, the amplitude forthe particle to be detected at the point Ois now given by the sum of three amplitudes,
the amplitude for the particle to propagate from the source Sthrough the hole A
1and
then onward to the point O, the amplitude for the particle to propagate from the source S
through the hole A2and then onward to the point O, and the amplitude for the particle to
propagate from the source Sthrough the hole A3and then onward to the point O.”
The professor was just about ready to continue when Feynman interjected again, “What
if I drill a fourth and a fifth hole in the screen?” Now the professor is visibly losing his
8 | I. Motivation and Foundation
SOA1
A2
Figure I.2.1
patience: “All right, wise guy, I think it is obvious to the whole class that we just sum over
all the holes.”
T o make what the professor said precise, denote the amplitude for the particle to
propagate from the source Sthrough the hole Aiand then onward to the point Oas
A(S→Ai→O). Then the amplitude for the particle to be detected at the point Ois
A(detected at O)=/summationdisplay
iA(S→Ai→O) (1)
But Feynman persisted, “What if we now add another screen (fig. I.2.2) with some holes
drilled in it?” The professor was really losing his patience: “Look, can’t you see that youjust take the amplitude to go from the source Sto the hole A
iin the first screen, then to
the hole Bjin the second screen, then to the detector at O, and then sum over all iandj?”
Feynman continued to pester, “What if I put in a third screen, a fourth screen, eh? What
if I put in a screen and drill an infinite number of holes in it so that the screen is no longerthere?” The professor sighed, “Let’s move on; there is a lot of material to cover in thiscourse.”
SOA1
A2
A3B1
B2
B3
B4
Figure I.2.2
I.2. Path Integral Formulation | 9
SO
Figure I.2.3
But dear reader, surely you see what that wise guy Feynman was driving at. I especially
enjoy his observation that if you put in a screen and drill an infinite number of holes in it,then that screen is not really there. Very Zen! What Feynman showed is that even if therewere just empty space between the source and the detector, the amplitude for the particleto propagate from the source to the detector is the sum of the amplitudes for the particle togo through each one of the holes in each one of the (nonexistent) screens. In other words,we have to sum over the amplitude for the particle to propagate from the source to thedetector following all possible paths between the source and the detector (fig. I.2.3).
A(particle to go from StoOin time T)=
/summationdisplay
(paths )A/parenleftbig
particle to go from StoOin time Tfollowing a particular path/parenrightbig
(2)
Now the mathematically rigorous will surely get anxious over how/summationtext
(paths )is to be
defined. Feynman followed Newton and Leibniz: T ake a path (fig. I.2.4), approximate itby straight line segments, and let the segments go to zero. You can see that this is just likefilling up a space with screens spaced infinitesimally close to each other, with an infinitenumber of holes drilled in each screen.
Fine, but how to construct the amplitude A(particle to go from StoOin time Tfollowing
a particular path)? Well, we can use the unitarity of quantum mechanics: If we know theamplitude for each infinitesimal segment, then we just multiply them together to get theamplitude of the whole path.
S
O
Figure I.2.4
10 | I. Motivation and Foundation
In quantum mechanics, the amplitude to propagate from a point qIto a point qFin
timeTis governed by the unitary operator e−iHT, where His the Hamiltonian. More
precisely, denoting by |q/angbracketrightthe state in which the particle is at q, the amplitude in question
is just /angbracketleftqF|e−iHT|qI/angbracketright. Here we are using the Dirac bra and ket notation. Of course,
philosophically, you can argue that to say the amplitude is /angbracketleftqF|e−iHT|qI/angbracketrightamounts to a
postulate and a definition of H. It is then up to experimentalists to discover that His
hermitean, has the form of the classical Hamiltonian, et cetera.
Indeed, the whole path integral formalism could be written down mathematically start-
ing with the quantity /angbracketleftqF|e−iHT|qI/angbracketright, without any of Feynman’s jive about screens with an
infinite number of holes. Many physicists would prefer a mathematical treatment withoutthe talk. As a matter of fact, the path integral formalism was invented by Dirac preciselyin this way, long before Feynman.
1
A necessary word about notation even though it interrupts the narrative flow: We denote
the coordinates transverse to the axis connecting the source to the detector by q, rather
thanx, for a reason which will emerge in a later chapter. For notational simplicity, we will
think of qas 1-dimensional and suppress the coordinate along the axis connecting the
source to the detector.
Dirac’s formulation
Let us divide the time TintoNsegments each lasting δt=T/N . Then we write
/angbracketleftqF|e−iHT|qI/angbracketright=/angbracketleftqF|e−iHδte−iHδt ...e−iHδt|qI/angbracketright
Our states are normalized by /angbracketleftq/prime|q/angbracketright=δ(q/prime−q)withδthe Dirac delta function. (Recall
thatδis defined by δ(q)=/integraltext∞
−∞(dp/2π)eipqand/integraltext
dqδ(q) =1. See appendix 1.) Now use
the fact that |q/angbracketrightforms a complete set of states so that/integraltext
dq|q/angbracketright/angbracketleftq|= 1. T o see that the
normalization is correct, multiply on the left by /angbracketleftq/prime/prime|and on the right by |q/prime/angbracketright, thus obtaining/integraltext
dqδ(q/prime/prime−q)δ(q −q/prime)=δ(q/prime/prime−q/prime). Insert 1 between all these factors of e−iHδtand write
/angbracketleftqF|e−iHT|qI/angbracketright
=(N−1/productdisplay
j=1/integraldisplay
dqj)/angbracketleftqF|e−iHδt|qN−1/angbracketright/angbracketleftqN−1|e−iHδt|qN−2/angbracketright.../angbracketleftq2|e−iHδt|q1/angbracketright/angbracketleftq 1|e−iHδt|qI/angbracketright (3)
Focus on an individual factor /angbracketleftqj+1|e−iHδt|qj/angbracketright. Let us take the baby step of first eval-
uating it just for the free-particle case in which H=ˆp2/2m. The hat on ˆpreminds us
that it is an operator. Denote by |p/angbracketrightthe eigenstate of ˆp, namely ˆp|p/angbracketright=p|p/angbracketright. Do you re-
member from your course in quantum mechanics that /angbracketleftq|p/angbracketright=eipq? Sure you do. This
1For the true history of the path integral, see p. xv of my introduction to R. P. Feynman, QED: The Strange
Theory of Light and Matter.
I.2. Path Integral Formulation | 11
just says that the momentum eigenstate is a plane wave in the coordinate representa-
tion. (The normalization is such that/integraltext
(dp/2π)|p/angbracketright/angbracketleftp|= 1. Again, to see that the nor-
malization is correct, multiply on the left by /angbracketleftq/prime|and on the right by |q/angbracketright, thus obtaining/integraltext
(dp/2π)eip(q/prime−q)=δ(q/prime−q).) So again inserting a complete set of states, we write
/angbracketleftqj+1|e−iδt( ˆp2/2m)|qj/angbracketright=/integraldisplaydp
2π/angbracketleftqj+1|e−iδt( ˆp2/2m)|p/angbracketright/angbracketleftp|qj/angbracketright
=/integraldisplaydp
2πe−iδt(p2/2m)/angbracketleftqj+1|p/angbracketright/angbracketleftp|qj/angbracketright
=/integraldisplaydp
2πe−iδt(p2/2m)eip(qj+1−qj)
Note that we removed the hat from the momentum operator in the exponential: Since the
momentum operator is acting on an eigenstate, it can be replaced by its eigenvalue. Also,we are evidently working in the Heisenberg picture.
The integral over pis known as a Gaussian integral, with which you may already be
familiar. If not, turn to appendix 2 to this chapter.
Doing the integral over p, we get (using (21))
/angbracketleftqj+1|e−iδt( ˆp2/2m)|qj/angbracketright=/parenleftbigg−im
2πδt/parenrightbigg1
2
e[im(qj+1−qj)2]/2δt
=/parenleftbigg−im
2πδt/parenrightbigg1
2
eiδt(m/2 )[(qj+1−qj)/δt ]2
Putting this into (3) yields
/angbracketleftqF|e−iHT|qI/angbracketright=/parenleftbigg−im
2πδt/parenrightbiggN
2/parenleftBiggN−1/productdisplay
k=1/integraldisplay
dqk/parenrightBigg
eiδt(m/2 )/Sigma1N−1
j=0[(qj+1−qj)/δt ]2
withq0≡qIandqN≡qF.
We can now go to the continuum limit δt→0. Newton and Leibniz taught us to replace
[(qj+1−qj)/δt ]2by˙q2, andδt/summationtextN−1
j=0by/integraltextT
0dt. Finally, we define the integral over paths
as
/integraldisplay
Dq(t) =lim
N→∞/parenleftbigg−im
2πδt/parenrightbiggN
2/parenleftBiggN−1/productdisplay
k=1/integraldisplay
dqk/parenrightBigg
We thus obtain the path integral representation
/angbracketleftqF|e−iHT|qI/angbracketright=/integraldisplay
Dq(t) ei/integraltextT
0dt1
2m˙q2
(4)
This fundamental result tells us that to obtain /angbracketleftqF|e−iHT|qI/angbracketrightwe simply integrate over
all possible paths q(t) such that q(0)=qIandq(T)=qF.
As an exercise you should convince yourself that had we started with the Hamiltonian
for a particle in a potential H=ˆp2/2m+V(ˆq)(again the hat on ˆqindicates an operator)
the final result would have been
/angbracketleftqF|e−iHT|qI/angbracketright=/integraldisplay
Dq(t) ei/integraltextT
0dt[1
2m˙q2−V( q) ](5)
12 | I. Motivation and Foundation
We recognize the quantity1
2m˙q2−V( q) as just the Lagrangian L(˙q,q). The Lagrangian
has emerged naturally from the Hamiltonian! In general, we have
/angbracketleftqF|e−iHT|qI/angbracketright=/integraldisplay
Dq(t) ei/integraltextT
0dtL(˙q,q)(6)
T o avoid potential confusion, let me be clear that tappears as an integration variable in
the exponential on the right-hand side. The appearance of tin the path integral measure
Dq(t) is simply to remind us that qis a function of t(as if we need reminding). Indeed,
this measure will often be abbreviated to Dq. You might recall that/integraltextT
0dtL(˙q,q)is called
the action S(q) in classical mechanics. The action Sis a functional of the function q(t) .
Often, instead of specifying that the particle starts at an initial position qIand ends at
a final position qF, we prefer to specify that the particle starts in some initial state Iand
ends in some final state F. Then we are interested in calculating /angbracketleftF|e−iHT|I/angbracketright, which upon
inserting complete sets of states can be written as
/integraldisplay
dqF/integraldisplay
dqI/angbracketleftF|qF/angbracketright/angbracketleftqF|e−iHT|qI/angbracketright/angbracketleftqI|I/angbracketright,
which mixing Schr ¨odinger and Dirac notation we can write as
/integraldisplay
dqF/integraldisplay
dqI/Psi1F(qF)∗/angbracketleftqF|e−iHT|qI/angbracketright/Psi1I(qI).
In most cases we are interested in taking |I/angbracketrightand|F/angbracketrightas the ground state, which we will
denote by |0/angbracketright. It is conventional to give the amplitude /angbracketleft0|e−iHT|0/angbracketrightthe name Z.
At the level of mathematical rigor we are working with, we count on the path integral
/integraltext
Dq(t) ei/integraltextT
0dt[1
2m˙q2−V( q) ]to converge because the oscillatory phase factors from different
paths tend to cancel out. It is somewhat more rigorous to perform a so-called Wick rotationto Euclidean time. This amounts to substituting t→−itand rotating the integration
contour in the complex tplane so that the integral becomes
Z=/integraldisplay
Dq(t) e−/integraltextT
0dt[1
2m˙q2+V( q) ], (7)
known as the Euclidean path integral. As is done in appendix 2 to this chapter with ordinary
integrals we will always assume that we can make this type of substitution with impunity.
The classical world emerges
One particularly nice feature of the path integral formalism is that the classical limit of
quantum mechanics can be recovered easily. We simply restore Planck’s constant /planckover2piin (6):
/angbracketleftqF|e−(i//planckover2pi )HT|qI/angbracketright=/integraldisplay
Dq(t) e(i//planckover2pi)/integraltextT
0dtL(˙q,q)
and take the /planckover2pi→0 limit. Applying the stationary phase or steepest descent method (if you
don’t know it see appendix 3 to this chapter) we obtain e(i//planckover2pi)/integraltextT
0dtL(˙qc,qc), where qc(t)is
the “classical path” determined by solving the Euler-Lagrange equation (d/dt)(δL/δ ˙q)−
(δL/δq) =0 with appropriate boundary conditions.
I.2. Path Integral Formulation | 13
Appendix 1
For your convenience, I include a concise review of the Dirac delta function here. Let us define a function dK(x)by
dK(x)≡/integraldisplayK
2
−K
2dk
2πeikx=1
πxsinKx
2(8)
for arbitrary real values of x. We see that for large Kthe even function dK(x) is sharply peaked at the origin x=0,
reaching a value of K/2πat the origin, crossing zero at x=2π/K , and then oscillating with ever decreasing
amplitude. Furthermore,
/integraldisplay∞
−∞dx dK(x)=2
π/integraldisplay∞
0dx
xsinKx
2=2
π/integraldisplay∞
0dy
ysiny=1 (9)
The Dirac delta function is defined by δ(x)=limK→∞dK(x). Heuristically, it could be thought of as an
infinitely sharp spike located at x=0 such that the area under the spike is equal to 1. Thus for a function s(x)
well-behaved around x=awe have
/integraldisplay∞
−∞dx δ(x −a)s(x) =s(a) (10)
(By the way, for what it is worth, mathematicians call the delta function a “distribution,” not a function.)
Our derivation also yields an integral representation for the delta function that we will use repeatedly in this
text:
δ(x)=/integraldisplay∞
−∞dk
2πeikx(11)
We will often use the identity
/integraldisplay∞
−∞dx δ(f (x))s(x) =/summationdisplay
is(xi)
|f/prime(xi)|(12)
where xidenotes the zeroes of f( x) (in other words, f( xi)=0 and f/prime(xi)=df (xi)/dx .) T o prove this, first show
that/integraltext∞
−∞dx δ(bx)s(x) =/integraltext∞
−∞dxδ(x)
|b|s(x)=s(0)/|b|. The factor of 1 /bfollows from dimensional analysis. (T o
see the need for the absolute value, simply note that δ(bx) is a positive function. Alternatively, change integration
variable to y=bx: forbnegative we have to flip the integration limits.) T o obtain (12), expand around each of
the zeroes of f( x) .
Another useful identity (understood in the limit in which the positive infinitesimal εtends to zero) is
1
x+iε=P1
x−iπδ(x) (13)
T o see this, simply write 1 /(x+iε)=x/(x2+ε2)−iε/(x2+ε2), and then note that ε/(x2+ε2)as a function
ofxis sharply spiked around x=0 and that its integral from −∞ to∞is equal to π. Thus we have another
representation of the Dirac delta function:
δ(x)=1
πε
x2+ε2(14)
Meanwhile, the principal value integral is defined by
/integraldisplay
dxP1
xf( x)=lim
ε→0/integraldisplay
dxx
x2+ε2f( x) (15)
14 | I. Motivation and Foundation
Appendix 2
I will now show you how to do the integral G≡/integraltext+∞
−∞dxe−1
2x2. The trick is to square the integral, call the dummy
integration variable in one of the integrals y, and then pass to polar coordinates:
G2=/integraldisplay+∞
−∞dx e−1
2x2/integraldisplay+∞
−∞dy e−1
2y2
=2π/integraldisplay+∞
0dr re−1
2r2
=2π/integraldisplay+∞
0dw e−w=2π
Thus, we obtain
/integraldisplay+∞
−∞dx e−1
2x2=√
2π (16)
Believe it or not, a significant fraction of the theoretical physics literature consists of varying and elaborating
this basic Gaussian integral. The simplest extension is almost immediate:
/integraldisplay+∞
−∞dx e−1
2ax2=/parenleftbigg2π
a/parenrightbigg1
2
(17)
as can be seen by scaling x→x/√a.
Acting on this repeatedly with −2(d/da) we obtain
/angbracketleftx2n/angbracketright≡/integraltext+∞
−∞dx e−1
2ax2x2n
/integraltext+∞
−∞dx e−1
2ax2=1
an(2n−1)(2n−3)...5.3.1 (18)
The factor 1 /anfollows from dimensional analysis. T o remember the factor (2n−1)!!≡(2n−1)(2n−3)...5.
3.1 imagine 2 npoints and connect them in pairs. The first point can be connected to one of (2n−1)points, the
second point can now be connected to one of the remaining (2n−3)points, and so on. This clever observation,
due to Gian Carlo Wick, is known as Wick’s theorem in the field theory literature. Incidentally, field theorists usethe following graphical mnemonic in calculating, for example, /angbracketleftx
6/angbracketright: Write /angbracketleftx6/angbracketrightas/angbracketleftxxxxxx/angbracketright and connect the x’s,
for example
〈〉xxxxxx
The pattern of connection is known as a Wick contraction. In this simple example, since the six x’s are identical,
any one of the distinct Wick contractions gives the same value a−3and the final result for /angbracketleftx6/angbracketrightis just a−3times
the number of distinct Wick contractions, namely 5 .3.1=15. We will soon come to a less trivial example, with
distinct x’s, in which case distinct Wick contraction gives distinct values.
An important variant is the integral
/integraldisplay+∞
−∞dx e−1
2ax2+Jx=/parenleftbigg2π
a/parenrightbigg1
2
eJ2/2a(19)
T o see this, take the expression in the exponent and “complete the square”: −ax2/2+Jx=−(a/2)(x2−
2Jx/a) =−(a/2)(x−J/a)2+J2/2a. The xintegral can now be done by shifting x→x+J/a , giving the
factor of (2π/a)1
2. Check that we can also obtain (18) by differentiating with respect to Jrepeatedly and then
setting J=0.
Another important variant is obtained by replacing JbyiJ:
/integraldisplay+∞
−∞dx e−1
2ax2+iJx=/parenleftbigg2π
a/parenrightbigg1
2
e−J2/2a(20)
I.2. Path Integral Formulation | 15
T o get yet another variant, replace aby−ia :
/integraldisplay+∞
−∞dx e1
2iax2+iJx=/parenleftbigg2πi
a/parenrightbigg1
2
e−iJ2/2a(21)
Let us promote ato a real symmetric NbyNmatrix Aijandxto a vector xi(i,j=1,... ,N). Then (19)
generalizes to
/integraldisplay+∞
−∞/integraldisplay+∞
−∞.../integraldisplay+∞
−∞dx1dx2...dxNe−1
2x.A.x+J.x=/parenleftbigg(2π)N
det[A]/parenrightbigg1
2
e1
2J.A−1.J(22)
where x.A.x=xiAijxjandJ.x=Jixi(with repeated indices summed.)
T o derive this important relation, diagonalize Aby an orthogonal transformation Oso that A=O−1.D.O,
where Dis a diagonal matrix. Call yi=Oijxj. In other words, we rotate the coordinates in the N-dimensional
Euclidean space we are integrating over. The expression in the exponential in the integrand then becomes
−1
2y.D.y+(OJ) .y. Using/integraltext+∞
−∞.../integraltext+∞
−∞dx1...dxN=/integraltext+∞
−∞.../integraltext+∞
−∞dy1...dyN, we factorize the left-hand
side of (22) into a product of Nintegrals, each of the form/integraltext+∞
−∞dyie−1
2Diiy2
i+(OJ) iyi. Plugging into (19) we
obtain the right hand side of (22), since (OJ) .D−1.(OJ)=J.O−1D−1O.J=J.A−1.J(where we use the
orthogonality of O). (T o make sure you got it, try this explicitly for N=2.)
Putting in some i’s (A→−iA,J→iJ), we find the generalization of (22)
/integraldisplay+∞
−∞/integraldisplay+∞
−∞.../integraldisplay+∞
−∞dx1dx2...dxNe(i/2)x.A.x+iJ.x
=/parenleftbigg(2πi)N
det[A]/parenrightbigg1
2
e−(i/2)J.A−1.J(23)
The generalization of (18) is also easy to obtain. Differentiate (22) ptimes with respect to Ji,Jj,...Jk, and
Jl, and then set J=0. For example, for p=1 the integrand in (22) becomes e−1
2x.A.xxiand since the integrand
is now odd in xithe integral vanishes. For p=2 the integrand becomes e−1
2x.A.x(xixj), while on the right hand
side we bring down A−1
ij. Rearranging and eliminating det[ A] (by setting J=0 in (22)), we obtain
/angbracketleftxixj/angbracketright=/integraltext+∞
−∞/integraltext+∞
−∞.../integraltext+∞
−∞dx1dx2...dxNe−1
2x.A.xxixj
/integraltext+∞
−∞/integraltext+∞
−∞.../integraltext+∞
−∞dx1dx2...dxNe−1
2x.A.x=A−1
ij
Just do it. Doing it is easier than explaining how to do it. Then do it for p=3 and 4. You will see immediately how
your result generalizes. When the set of indices i,j,... ,k,lcontains an odd number of elements, /angbracketleftxixj...xkxl/angbracketright
vanishes trivially. When the set of indices i,j,... ,k,lcontains an even number of elements, we have
/angbracketleftxixj...xkxl/angbracketright=/summationdisplay
Wick(A−1)ab...(A−1)cd (24)
where we have defined
/angbracketleftxixj...xkxl/angbracketright
=/integraltext+∞
−∞/integraltext+∞
−∞.../integraltext+∞
−∞dx1dx2...dxNe−1
2x.A.xxixj...xkxl
/integraltext+∞
−∞/integraltext+∞
−∞.../integraltext+∞
−∞dx1dx2...dxNe−1
2x.A.x(25)
and where the set of indices {a,b,... ,c,d}represent a permutation of {i,j,... ,k,l}. The sum in (24) is over
all such permutations or Wick contractions.
For example,
/angbracketleftxixjxkxl/angbracketright=(A−1)ij(A−1)kl+(A−1)il(A−1)jk+(A−1)ik(A−1)jl (26)
(Recall that A, and thus A−1, is symmetric.) As in the simple case when xdoes not carry any index, we could
connect the x’s in/angbracketleftxixjxkxl/angbracketrightin pairs (Wick contraction) and write a factor (A−1)abif we connect xatoxb.
Notice that since /angbracketleftxixj/angbracketright=(A−1)ijthe right hand side of (24) can also be written in terms of objects like /angbracketleftxixj/angbracketright.
Thus, /angbracketleftxixjxkxl/angbracketright=/angbracketleftxixj/angbracketright/angbracketleftxkxl/angbracketright+/angbracketleftxixl/angbracketright/angbracketleftxjxk/angbracketright+/angbracketleftxixk/angbracketright/angbracketleftxjxl/angbracketright.
16 | I. Motivation and Foundation
Please work out /angbracketleftxixjxkxlxmxn/angbracketright; you will become an expert on Wick contractions. Of course, (24) reduces to
(18) for N=1.
Perhaps you are like me and do not like to memorize anything, but some of these formulas might be worth
memorizing as they appear again and again in theoretical physics (and in this book).
Appendix 3
T o do an exponential integral of the form I=/integraltext+∞
−∞dqe−(1//planckover2pi)f (q)we often have to resort to the steepest-descent
approximation, which I will now review for your convenience. In the limit of /planckover2pismall, the integral is dominated
by the minimum of f( q) . Expanding f( q)=f( a)+1
2f/prime/prime(a)(q−a)2+O[(q−a)3] and applying (17) we obtain
I=e−(1//planckover2pi)f (a)/parenleftbigg2π/planckover2pi
f/prime/prime(a)/parenrightbigg1
2
e−O(/planckover2pi1
2)(27)
Forf( q) a function of many variables q1,..., qNand with a minimum at qj=aj, we generalize immediately
to
I=e−(1//planckover2pi)f (a)/parenleftbigg(2π/planckover2pi)N
detf/prime/prime(a)/parenrightbigg1
2
e−O(/planckover2pi1
2)(28)
Heref/prime/prime(a) denotes the NbyNmatrix with entries [f/prime/prime(a)]ij≡(∂2f/∂qi∂qj)|q=a. In many situations, we do
not even need the factor involving the determinant in (28). If you can derive (28) you are well on your way tobecoming a quantum field theorist!
Exercises
I.2.1 Verify (5).
I.2.2 Derive (24).
I.3 From Mattress to Field
The mattress in the continuum limit
The path integral representation
Z≡/angbracketleft 0|e−iHT|0/angbracketright=/integraldisplay
Dq(t) ei/integraltextT
0dt[1
2m˙q2−V( q) ](1)
(we suppress the factor /angbracketleft0|qf/angbracketright/angbracketleftqI|0/angbracketright; we will come back to this issue later in this chapter)
which we derived for the quantum mechanics of a single particle, can be generalized almostimmediately to the case of Nparticles with the Hamiltonian
H=/summationdisplay
a1
2maˆp2
a+V(ˆq1,ˆq2,... ,ˆqN). (2)
We simply keep track mentally of the position of the particles qawitha=1, 2, ... ,N.
Going through the same steps as before, we obtain
Z≡/angbracketleft 0|e−iHT|0/angbracketright=/integraldisplay
Dq(t) eiS(q)(3)
with the action
S(q)=/integraldisplayT
0dt/parenleftBig/summationdisplay
a1
2ma˙q2
a−V[q1,q2,... ,qN]/parenrightBig
.
The potential energy V( q1,q2,...,qN)now includes interaction energy between particles,
namely terms of the form v(qa−qb), as well as the energy due to an external potential,
namely terms of the form w(qa). In particular, let us now write the path integral description
of the quantum dynamics of the mattress described in chapter I.1, with the potential
V( q 1,q2,..., qN)=/summationdisplay
ab1
2kab(qa−qb)2+...
We are now just a short hop and skip away from a quantum field theory! Suppose we
are only interested in phenomena on length scales much greater than the lattice spacingl(see fig. I.1.1). Mathematically, we take the continuum limit l→0. In this limit, we can
18 | I. Motivation and Foundation
replace the label aon the particles by a two-dimensional position vector /vectorx, and so we write
q(t,/vectorx)instead of qa(t). It is traditional to replace the Latin letter qby the Greek letter ϕ.
The function ϕ(t,/vectorx)is called a field.
The kinetic energy/summationtext
a1
2ma˙q2
anow becomes/integraltext
d2x1
2σ(∂ϕ/∂t)2. We replace/summationtext
aby/integraltext
d2x/l2and denote the mass per unit area ma/l2byσ. We take all the ma’s to be equal;
otherwise σwould be a function of /vectorx, the system would be inhomogeneous, and we would
have a hard time writing down a Lorentz-invariant action (see later).
We next focus on the first term in V. Assume for simplicity that kabconnect only nearest
neighbors on the lattice. For nearest-neighbor pairs (qa−qb)2/similarequall2(∂ϕ/∂x)2+... in the
continuum limit; the derivative is obviously taken in the direction that joins the lattice sitesaandb.
Putting it together then, we have
S(q)→S(ϕ)≡/integraldisplayT
0dt/integraldisplay
d2xL(ϕ)
=/integraldisplayT
0dt/integraldisplay
d2x1
2/braceleftBigg
σ/parenleftbigg∂ϕ
∂t/parenrightbigg2
−ρ/bracketleftBigg/parenleftbigg∂ϕ
∂x/parenrightbigg2
+/parenleftbigg∂ϕ
∂y/parenrightbigg2/bracketrightBigg
−τϕ2−ςϕ4+.../bracerightBigg
(4)
where the parameter ρis determined by kabandl. The precise relations do not concern us.
Henceforth in this book, we will take the T→∞ limit so that we can integrate over all
of spacetime in (4).
We can clean up a bit by writing ρ=σc2and scaling ϕ→ϕ/√σ, so that the combination
(∂ϕ/∂t)2−c2[(∂ϕ/∂x)2+(∂ϕ/∂y)2] appears in the Lagrangian. The parameter cevidently
has the dimension of a velocity and defines the phase velocity of the waves on our mattress.
We started with a mattress for pedagogical reasons. Of course nobody believes that
the fields observed in Nature, such as the meson field or the photon field, are actuallyconstructed of point masses tied together with springs. The modern view, which I will callLandau-Ginzburg, is that we start with the desired symmetry, say Lorentz invariance if wewant to do particle physics, decide on the fields we want by specifying how they transformunder the symmetry (in this case we decided on a scalar field ϕ), and then write down
the action involving no more than two time derivatives (because we don’t know how toquantize actions with more than two time derivatives).
We end up with a Lorentz-invariant action (setting c=1)
S=/integraldisplay
ddx/bracketleftbigg1
2(∂ϕ)2−1
2m2ϕ2−g
3!ϕ3−λ
4!ϕ4+.../bracketrightbigg
(5)
where various numerical factors are put in for later convenience. The relativistic nota-
tion(∂ϕ)2≡∂μϕ∂μϕ=(∂ϕ/∂t)2−(∂ϕ/∂x)2−(∂ϕ/∂y)2was explained in the note on
convention. The dimension of spacetime, d, clearly can be any integer, even though in
our mattress model it was actually 3. We often write d=D+1 and speak of a ( D+1)-
dimensional spacetime.
We see here the power of imposing a symmetry. Lorentz invariance together with the
insistence that the Lagrangian involve only at most two powers of ∂/∂t immediately tells us
I.3. From Mattress to Field | 19
that the Lagrangian can only have the form1L=1
2(∂ϕ)2−V( ϕ) withVsome function of
ϕ. For simplicity, we now restrict Vto be a polynomial in ϕ, although much of the present
discussion will not depend on this restriction. We will have a great deal more to say aboutsymmetry later. Here we note that, for example, we could insist that physics is symmetricunder ϕ→−ϕ, in which case V( ϕ) would have to be an even polynomial.
Now that you know what a quantum field theory is, you realize why I used the letter q
to label the position of the particle in the previous chapter and not the more common /vectorx.
In quantum field theory, /vectorxis a label, not a dynamical variable. The /vectorxappearing in ϕ(t,/vectorx)
corresponds to the label ainq
a(t)in quantum mechanics. The dynamical variable in field
theory is not position, but the field ϕ. The variable /vectorxsimply specifies which field variable we
are talking about. I belabor this point because upon first exposure to quantum field theorysome students, used to thinking of /vectorxas a dynamical operator in quantum mechanics, are
confused by its role here.
In summary, we have the table
q→ϕ
a→/vectorx
(6)
qa(t)→ϕ(t,/vectorx)=ϕ(x)
/summationtext
a→/integraltext
dDx
Thus we finally have the path integral defining a scalar field theory in d=(D+1)dimen-
sional spacetime:
Z=/integraldisplay
Dϕei/integraltext
ddx(1
2(∂ϕ)2−V( ϕ) )(7)
Note that a (0 +1)-dimensional quantum field theory is just quantum mechanics.
The classical limit
As I have already remarked, the path integral formalism is particularly convenient for
taking the classical limit. Remembering that Planck’s constant /planckover2pihas the dimension of
energy multiplied by time, we see that it appears in the unitary evolution operator e(−i//planckover2pi)HT.
T racing through the derivation of the path integral, we see that we simply divide the overallfactor iby/planckover2pito get
Z=/integraldisplay
Dϕe(i//planckover2pi)/integraltext
d4xL(ϕ)(8)
1Strictly speaking, a term of the form U(ϕ)(∂ϕ)2is also possible. In quantum mechanics, a term such as
U(q)(dq/dt)2in the Lagrangian would describe a particle whose mass depends on position. We will not consider
such “nasty” terms until much later.
20 | I. Motivation and Foundation
In the limit /planckover2pimuch smaller than the relevant action we are considering, we can evaluate the
path integral using the stationary phase (or steepest descent) approximation, as I explainedin the previous chapter in the context of quantum mechanics. We simply determine theextremum of/integraltext
d
4xL(ϕ). According to the usual Euler-Lagrange variational procedure, this
leads to the equation
∂μδL
δ(∂μϕ)−δL
δϕ=0 (9)
We thus recover the classical field equation, exactly as we should, which in our scalar field
theory reads
(∂2+m2)ϕ(x) +g
2ϕ(x)2+λ
6ϕ(x)3+...=0 (10)
The vacuum
In the point particle quantum mechanics discussed in chapter I.2 we wrote the path
integral for /angbracketleftF|e−iHT|I/angbracketright, with some initial and final state, which we can choose at our
pleasure. A convenient and particularly natural choice would be to take |I/angbracketright=|F/angbracketrightto be
the ground state. In quantum field theory what should we choose for the initial and finalstates? A standard choice for the initial and final states is the ground state or the vacuumstate of the system, denoted by |0/angbracketright, in which, speaking colloquially, nothing is happening.
In other words, we would calculate the quantum transition amplitude from the vacuum tothe vacuum, which would enable us to determine the energy of the ground state. But thisis not a particularly interesting quantity, because in quantum field theory we would like tomeasure all energies relative to the vacuum and so, by convention, would set the energyof the vacuum to zero (possibly by having to subtract a constant from the Lagrangian).Incidentally, the vacuum in quantum field theory is a stormy sea of quantum fluctuations,but for this initial pass at quantum field theory, we will not examine it in any detail. Wewill certainly come back to the vacuum in later chapters.
Disturbing the vacuum
We might enjoy doing something more exciting than watching a boiling sea of quantum
fluctuations. We might want to disturb the vacuum. Somewhere in space, at some instant
in time, we would like to create a particle, watch it propagate for a while, and then annihilateit somewhere else in space, at some later instant in time. In other words, we want to setup a source and a sink (sometimes referred to collectively as sources) at which particlescan be created and annihilated.
T o see how to do this, let us go back to the mattress. Bounce up and down on it to create
some excitations. Obviously, pushing on the mass labeled by ain the mattress corresponds
to adding a term such as J
a(t)qato the potential V( q1,q2,... ,qN). More generally,
I.3. From Mattress to Field | 21
JJJ
?
t
x
Figure I.3.1
we can add/summationtext
aJa(t)qa. When we go to field theory this added term gets promoted to/integraltext
dDxJ(x)ϕ(x) in the field theory Lagrangian, according to the promotion table (6).
This so-called source function J(t,/vectorx)describes how the mattress is being disturbed.
We can choose whatever function we like, corresponding to our freedom to push on themattress wherever and whenever we like. In particular, J(x) can vanish everywhere in
spacetime except in some localized regions.
By bouncing up and down on the mattress we can get wave packets going off here and
there (fig. I.3.1). This corresponds precisely to sources (and sinks) for particles. Thus, wereally want the path integral
Z=/integraldisplay
Dϕei/integraltext
d4x[1
2(∂ϕ)2−V( ϕ) +J(x)ϕ(x) ](11)
Free field theory
The functional integral in (11) is impossible to do except when
L(ϕ)=1
2[(∂ϕ)2−m2ϕ2] (12)
The corresponding theory is called the free or Gaussian theory. The equation of motion
(9) works out to be (∂2+m2)ϕ=0, known as the Klein-Gordon equation.2Being linear, it
can be solved immediately to give ϕ(/vectorx,t)=ei(ωt−/vectork./vectorx)with
ω2=/vectork2+m2(13)
2The Klein-Gordon equation was actually discovered by Schr ¨odinger before he found the equation that now
bears his name. Later, in 1926, it was written down independently by Klein, Gordon, Fock, Kudar, de Donder,
and Van Dungen.
22 | I. Motivation and Foundation
In the natural units we are using, /planckover2pi=1 and so frequency ωis the same as energy /planckover2piω
and wave vector /vectorkis the same as momentum /planckover2pi/vectork. Thus, we recognize (13) as the energy-
momentum relation for a particle of mass m, namely the sophisticate’s version of the
layperson’s E=mc2. We expect this field theory to describe a relativistic particle of mass m.
Let us now evaluate (11) in this special case:
Z=/integraldisplay
Dϕei/integraltext
d4x{1
2[(∂ϕ)2−m2ϕ2]+Jϕ}(14)
Integrating by parts under the/integraltext
d4xand not worrying about the possible contribution of
boundary terms at infinity (we implicitly assume that the fields we are integrating over falloff sufficiently rapidly), we write
Z=/integraldisplay
Dϕei/integraltext
d4x[−1
2ϕ(∂2+m2)ϕ+Jϕ ](15)
You will encounter functional integrals like this again and again in your study of field
theory. The trick is to imagine discretizing spacetime. You don’t actually have to do it:Just imagine doing it. Let me sketch how this goes. Replace the function ϕ(x) by the
vector ϕ
i=ϕ(ia) withian integer and athe lattice spacing. (For simplicity, I am writing
things as if we were in 1-dimensional spacetime. More generally, just let the index i
enumerate the lattice points in some way.) Then differential operators become matrices.For example, ∂ϕ(ia) →(1/a) (ϕ
i+1−ϕi)≡/summationtext
jMijϕj, with some appropriate matrix M.
Integrals become sums. For example,/integraltext
d4xJ(x)ϕ(x) →a4/summationtext
iJiϕi.
Now, lo and behold, the integral (15) is just the integral we did in (I.2.23)
/integraldisplay+∞
−∞/integraldisplay+∞
−∞.../integraldisplay+∞
−∞dq1dq2...dqNe(i/2)q.A.q+iJ.q
=/parenleftbigg(2πi)N
det[A]/parenrightbigg1
2
e−(i/2)J.A−1.J(16)
The role of Ain (16) is played in (15) by the differential operator −(∂2+m2). The defining
equation for the inverse, A.A−1=IorAijA−1
jk=δik, becomes in the continuum limit
−(∂2+m2)D(x−y)=δ(4)(x−y) (17)
We denote the continuum limit of A−1
jkbyD(x−y)(which we know must be a function
ofx−y, and not of xandyseparately, since no point in spacetime is special). Note that
in going from the lattice to the continuum Kronecker is replaced by Dirac. It is very usefulto be able to go back and forth mentally between the lattice and the continuum.
Our final result is
Z(J)=Ce−(i/2)/integraltext/integraltext
d4xd4yJ(x)D(x −y)J(y)≡CeiW(J)(18)
withD(x) determined by solving (17). The overall factor C, which corresponds to the overall
factor with the determinant in (16), does not depend on Jand, as will become clear in the
discussion to follow, is often of no interest to us. I will often omit writing Caltogether.
Clearly, C=Z(J=0)so that W(J) is defined by
Z(J)≡Z(J=0)eiW(J)(19)
I.3. From Mattress to Field | 23
Observe that
W(J) =−1
2/integraldisplay/integraldisplay
d4xd4yJ(x)D(x −y)J(y) (20)
is a simple quadratic functional of J. In contrast, Z(J) depends on arbitrarily high powers
ofJ. This fact will be of importance in chapter I.7.
Free propagator
The function D(x) , known as the propagator, plays an essential role in quantum field
theory. As the inverse of a differential operator it is clearly closely related to the Green’sfunction you encountered in a course on electromagnetism.
Physicists are sloppy about mathematical rigor, but even so, they have to be careful once
in a while to make sure that what they are doing actually makes sense. For the integralin (15) to converge for large ϕwe replace m
2→m2−iεso that the integrand contains a
factor e−ε/integraltext
d4xϕ2, where εis a positive infinitesimal we will let tend to zero.3
We can solve (17) easily by going to momentum space and multiplying together four
copies of the representation (I.2.11) of the Dirac delta function
δ(4)(x−y)=/integraldisplayd4k
(2π)4eik(x−y)(21)
The solution is
D(x−y)=/integraldisplayd4k
(2π)4eik(x−y)
k2−m2+iε(22)
which you can check by plugging into (17):
−(∂2+m2)D(x−y)=/integraldisplayd4k
(2π)4k2−m2
k2−m2+iεeik(x−y)=/integraldisplayd4k
(2π)4eik(x−y)=δ(4)(x−y)asε→0.
Note that the so-called iεprescription we just mentioned is essential; otherwise the
integral giving D(x) would hit a pole. The magnitude of εis not important as long as it is
infinitesimal, but the positive sign of εis crucial as we will see presently. (More on this in
chapter III.8.) Also, note that the sign of kin the exponential does not matter here by the
symmetry k→−k.
T o evaluate D(x) we first integrate over k0by the method of contours. Define ωk≡
+/radicalbig
/vectork2+m2with a plus sign. The integrand has two poles in the complex k0plane, at
±/radicalBig
ω2
k−iε, which in the ε→0 limit are equal to +ωk−iεand−ωk+iε. Thus for ε
positive, one pole is in the lower half-plane and the other in the upper half plane, and so
as we go along the real k0axis from −∞ to+∞ we do not run into the poles. The issue is
how to close the integration contour.
Forx0positive, the factor eik0x0is exponentially damped for k0in the upper half-plane.
Hence we should extend the integration contour extending from −∞ to+∞ on the real
3As is customary, εis treated as generic, so that εmultiplied by any positive number is still ε.
24 | I. Motivation and Foundation
axis to include the infinite semicircle in the upper half-plane, thus enclosing the pole at
−ωk+iεand giving −i/integraltextd3k
(2π)32ωke−i(ω kt−/vectork./vectorx). Again, note that we are free to flip the sign
of/vectork. Also, as is conventional, we use x0andtinterchangeably. (In view of some reader
confusion here in the first edition, I might add that I generally use x0withk0andtwith
ωk;k0is a variable that can take on either sign but ωkis a positive function of /vectork.)
Forx0negative, we do the opposite and close the contour in the lower half-plane, thus
picking up the pole at +ωk−iε. We now obtain −i/integraltext
(d3k/(2π)32ωk)e+i(ω kt−/vectork./vectorx).
Recall that the Heaviside (we will meet this great and aptly named physicist in chapter
IV .4) step function θ(t) is defined to be equal to 0 for t< 0 and equal to 1 for t> 0. As for
whatθ(0)should be, the answer is that since we are proud physicists and not nitpicking
mathematicians we will just wing it when the need arises. The step function allows us topackage our two integration results together as
D(x)=−i/integraldisplayd3k
(2π)32ωk[e−i(ω kt−/vectork./vectorx)θ(x0)+ei(ωkt−/vectork./vectorx)θ(−x0)] (23)
Physically, D(x) describes the amplitude for a disturbance in the field to propagate from
the origin to x. Lorentz invariance tells us that it is a function of x2and the sign of x0(since
these are the quantities that do not change under a Lorentz transformation). We thus expectdrastically different behavior depending on whether xis inside or outside the lightcone
defined by x
2=(x0)2−/vectorx2=0. Without evaluating the d3kintegral we can see roughly
how things go. Let us look at some cases.
In the future cone, x=(t,0)witht> 0,D(x)=−i/integraltext
(d3k/(2π)32ωk)e−iωkta superpo-
sition of plane waves and thus D(x) oscillates. In the past cone, x=(t,0)witht< 0,
D(x)=−i/integraltext
(d3k/(2π)32ωk)e+iωktoscillates with the opposite phase.
In contrast, for xspacelike rather than timelike, x0=0, we have, upon interpret-
ingθ(0)=1
2(the obvious choice; imagine smoothing out the step function), D(x)=
−i/integraltext
(d3k/(2π)32/radicalbig
/vectork2+m2)e−i/vectork./vectorx. The square root cut starting at ±im tells us that the
characteristic value of |/vectork|in the integral is of order m, leading to an exponential decay
∼e−m|/vectorx|, as we would expect. Classically, a particle cannot get outside the lightcone, but a
quantum field can “leak” out over a distance of order m−1by the Heisenberg uncertainty
principle.
Exercises
I.3.1 Verify that D(x) decays exponentially for spacelike separation.
I.3.2 Work out the propagator D(x) for a free field theory in (1+1)-dimensional spacetime and study the large
x1behavior for x0=0.
I.3.3 Show that the advanced propagator defined by
Dadv(x−y)=/integraldisplayd4k
(2π)4eik(x−y)
k2−m2−isgn(k0)ε
I.3. From Mattress to Field | 25
(where the sign function is defined by sgn (k0)=+ 1i fk0>0 and sgn (k0)=− 1i fk0<0) is nonzero only
ifx0>y0. In other words, it only propagates into the future. [Hint: both poles of the integrand are now
in the upper half of the k0-plane.] Incidentally, some authors prefer to write (k0−ie)2−/vectork2−m2instead
ofk2−m2−isgn(k0)εin the integrand. Similarly, show that the retarded propagator
Dret(x−y)=/integraldisplayd4k
(2π)4eik(x−y)
k2−m2+isgn(k0)ε
propagates into the past.
I.4 From Field to Particle to Force
From field to particle
In the previous chapter we obtained for the free theory
W(J) =−1
2/integraldisplay/integraldisplay
d4xd4yJ(x)D(x −y)J(y) (1)
which we now write in terms of the Fourier transform J(k)≡/integraltext
d4xe−ikxJ(x) :
W(J) =−1
2/integraldisplayd4k
(2π)4J(k)∗ 1
k2−m2+iεJ(k) (2)
[Note that J(k)∗=J(−k) forJ(x) real.]
We can jump up and down on the mattress any way we like. In other words, we can
choose any J(x) we want, and by exploiting this freedom of choice, we can extract a
remarkable amount of physics.
Consider J(x)=J1(x)+J2(x), where J1(x) andJ2(x) are concentrated in two local
regions 1 and 2 in spacetime (fig. I.4.1). Then W(J) contains four terms, of the form J∗
1J1,
J∗
2J2,J∗
1J2, andJ∗
2J1. Let us focus on the last two of these terms, one of which reads
W(J) =−1
2/integraldisplayd4k
(2π)4J2(k)∗ 1
k2−m2+iεJ1(k) (3)
We see that W(J) is large only if J1(x) andJ2(x) overlap significantly in their Fourier
transform and if in the region of overlap in momentum space k2−m2almost vanishes.
There is a “resonance type” spike at k2=m2, that is, if the energy-momentum relation
of a particle of mass mis satisfied. (We will use the language of the relativistic physicist,
writing “momentum space” for energy-momentum space, and lapse into nonrelativisticlanguage only when the context demands it, such as in “energy-momentum relation.”)
We thus interpret the physics contained in our simple field theory as follows: In region
1 in spacetime there exists a source that sends out a “disturbance in the field,” whichis later absorbed by a sink in region 2 in spacetime. Experimentalists choose to call this
I.4. From Field to Particle to Force | 27
J1tJ2
x
Figure I.4.1
disturbance in the field a particle of mass m. Our expectation based on the equation of
motion that the theory contains a particle of mass mis fulfilled.
A bit of jargon: When k2=m2,kis said to be on mass shell. Note, however, that in (3)
we integrate over all k, including values of kfar from the mass shell. For arbitrary k,i ti s
a linguistic convenience to say that a “virtual particle” of momentum kpropagates from
the source to the sink.
From particle to force
We can now go on to consider other possibilities for J(x) (which we will refer to generically
as sources), for example, J(x)=J1(x)+J2(x), where Ja(x)=δ(3)(/vectorx−/vectorxa). In other words,
J(x) is a sum of sources that are time-independent infinitely sharp spikes located at /vectorx1and
/vectorx2in space. (If you like more mathematical rigor than is offered here, you are welcome to
replace the delta function by lumpy functions peaking at /vectorxa. You would simply clutter up
the formulas without gaining much.) More picturesquely, we are describing two massive
lumps sitting at /vectorx1and/vectorx2on the mattress and not moving at all [no time dependence in
J(x) ].
What do the quantum fluctuations in the field ϕ, that is, the vibrations in the mattress,
do to the two lumps sitting on the mattress? If you expect an attraction between the twolumps, you are quite right.
As before, W(J) contains four terms. We neglect the “self-interaction” term J
1J1since
this contribution would be present in Wregardless of whether J2is present or not. We
want to study the interaction between the two “massive lumps” represented by J1andJ2.
Similarly we neglect J2J2.
28 | I. Motivation and Foundation
Plugging into (1) and doing the integral over d3xandd3ywe immediately obtain
W(J) =−/integraldisplay/integraldisplay
dx0dy0/integraldisplaydk0
2πeik0(x−y)0/integraldisplayd3k
(2π)3ei/vectork.(/vectorx1−/vectorx2)
k2−m2+iε(4)
(The factor 2 comes from the two terms J2J1andJ1J2. ) Integrating over y0we get a delta
function setting k0to zero (so that kis certainly not on mass shell, to throw the jargon
around a bit). Thus we are left with
W(J) =/parenleftbigg/integraldisplay
dx0/parenrightbigg/integraldisplayd3k
(2π)3ei/vectork.(/vectorx1−/vectorx2)
/vectork2+m2(5)
Note that the infinitesimal iεcan be dropped since the denominator /vectork2+m2is always
positive.
The factor (/integraltext
dx0)should have filled us with fear and trepidation: an integral over time,
it seems to be infinite. Fear not! Recall that in the path integral formalism Z=CeiW(J)
represents /angbracketleft0|e−iHT|0/angbracketright=e−iET, where Eis the energy due to the presence of the two
sources acting on each other. The factor (/integraltext
dx0)produces precisely the time interval T. All
is well. Setting iW=−iET we obtain from (5)
E=−/integraldisplayd3k
(2π)3ei/vectork.(/vectorx1−/vectorx2)
/vectork2+m2(6)
The integral is evaluated in an appendix. This energy is negative! The presence of two delta
function sources, at /vectorx1and/vectorx2, has lowered the energy. (Notice that for the two sources
infinitely far apart, we have, as we might expect, E=0: the infinitely rapidly oscillating
exponential kills the integral.) In other words, two like objects attract each other by virtueof their coupling to the field ϕ. We have derived our first physical result in quantum field
theory!
We identify Eas the potential energy between two static sources. Even without doing
the integral, we see by dimensional analysis that the characteristic distance beyond whichthe integral goes to zero is given by the inverse of the characteristic value of k, which is
m. Thus, we expect the attraction between the two sources to decrease rapidly to zero over
the distance 1 /m.
The range of the attractive force generated by the field ϕis determined inversely by the
massmof the particle described by the field. Got that?
The integral is done in the appendix to this chapter and gives
E=−1
4πre−mr(7)
The result is as we expected: The potential drops off exponentially over the distance scale
1/m. Obviously, dE/dr > 0: The two massive lumps sitting on the mattress can lower the
energy by getting closer to each other.
What we have derived was one of the most celebrated results in twentieth-century
physics. Yukawa proposed that the attraction between nucleons in the atomic nucleus isdue to their coupling to a field like the ϕfield described here. The known range of the
nuclear force enabled him to predict not only the existence of the particle associated with
I.4. From Field to Particle to Force | 29
this field, now called the πmeson1or the pion, but its mass as well. As you probably know,
the pion was eventually discovered with essentially the properties predicted by Yukawa.
Origin of force
That the exchange of a particle can produce a force was one of the most profound concep-tual advances in physics. We now associate a particle with each of the known forces: forexample, the photon with the electromagnetic force and the graviton with the gravitationalforce; the former is experimentally well established and while the latter has not yet beendetected experimentally hardly anyone doubts its existence. We will discuss the photon andthe graviton in the next chapter, but we can already answer a question smart high schoolstudents often ask: Why do Newton’s gravitational force and Coulomb’s electric force bothobey the 1 /r
2law?
We see from (7) that if the mass mof the mediating particle vanishes, the force produced
will obey the 1 /r2law. If you trace back over our derivation, you will see that this comes
from the fact that the Lagrangian density for the simplest field theory involves two powersof the spacetime derivative ∂(since any term involving one derivative such as ϕ∂ ϕ is not
Lorentz invariant). Indeed, the power dependence of the potential follows simply from
dimensional analysis:/integraltext
d
3k(ei/vectork./vectorx/k2)∼1/r.
Connected versus disconnected
We end with a couple of formal remarks of importance to us only in chapter I.7. First,
note that we might want to draw a small picture fig. (I.4.2) to represent the integrandJ(x)D(x −y)J(y) inW(J) : A disturbance propagates from ytox(or vice versa). In fact,
this is the beginning of Feynman diagrams! Second, recall that
Z(J)=Z(J=0)∞/summationdisplay
n=0[iW(J)]n
n!
For instance, the n=2 term in Z(J)/Z(J =0)is given by
1
2!/parenleftbigg
−i
2/parenrightbigg2/integraldisplay/integraldisplay/integraldisplay/integraldisplay
d4x1d4x2d4x3d4x4D(x 1−x2)
D(x 3−x4)J(x1)J(x2)J(x3)J(x4)
The integrand is graphically described in figure I.4.3. The process is said to be discon-
nected: The propagation from x1tox2and the propagation from x3tox4proceed inde-
pendently. We will come back to the difference between connected and disconnected inchapter I.7.
1The etymology behind this word is quite interesting (A. Zee, Fearful Symmetry : see pp. 169 and 335 to learn,
among other things, the French objection and the connection between meson and illusion).
30 | I. Motivation and Foundation
yx
Figure I.4.2
x1x2
x4
x3
Figure I.4.3
I.4. From Field to Particle to Force | 31
Appendix
Writing /vectorx≡(/vectorx1−/vectorx2)andu≡cosθwithθthe angle between /vectorkand/vectorx, we evaluate the integral in (6) in spherical
coordinates (with k=|/vectork|andr=|/vectorx|):
I≡1
(2π)2/integraldisplay∞
0dk k2/integraldisplay+1
−1dueikru
k2+m2=2i
(2π)2ir/integraldisplay∞
0dk ksinkr
k2+m2(8)
Since the integrand is even, we can extend the integral and write it as
1
2/integraldisplay∞
−∞dk ksinkr
k2+m2=1
2i/integraldisplay∞
−∞dk k1
k2+m2eikr.
Since ris positive, we can close the contour in the upper half-plane and pick up the pole at +im , obtaining
(1/2i)(2πi)(im/2 im)e−mr=(π/2)e−mr. Thus, I=(1/4πr)e−mr.
Exercise
I.4.1 Calculate the analog of the inverse square law in a (2 +1)-dimensional universe, and more generally in
a(D+1)-dimensional universe.
I.5 Coulomb and Newton: Repulsion and Attraction
Why like charges repel
We suggested that quantum field theory can explain both Newton’s gravitational force and
Coulomb’s electric force naturally. Between like objects Newton’s force is attractive whileCoulomb’s force is repulsive. Is quantum field theory “smart enough” to produce thisobservational fact, one of the most basic in our understanding of the physical universe?You bet!
We will first treat the quantum field theory of the electromagnetic field, known as
quantum electrodynamics or QED for short. In order to avoid complications at this stageassociated with gauge invariance (about which much more later) I will consider insteadthe field theory of a massive spin 1 meson, or vector meson. After all, experimentally all weknow is an upper bound on the photon mass, which although tiny is not mathematicallyzero. We can adopt a pragmatic attitude: Calculate with a photon mass mand set m=0a t
the end, and if the result does not blow up in our faces, we will presume that it is OK.
1
Recall Maxwell’s Lagrangian for electromagnetism L=−1
4FμνFμν, where Fμν≡∂μAν
−∂νAμwithAμ(x) the vector potential. You can see the reason for the important overall
minus sign in the Lagrangian by looking at the coefficient of (∂0Ai)2, which has to be
positive, just like the coefficient of (∂0ϕ)2in the Lagrangian for the scalar field. This says
simply that time variation should cost a positive amount of action.
I will now give the photon a small mass by changing the Lagrangian to L=−1
4FμνFμν
+1
2m2AμAμ+AμJμ. (The mass term is written in analogy to the mass term m2ϕ2in
the scalar field Lagrangian; we will see shortly that the sign is correct and that this term
indeed leads to a photon mass.) I have also added a source Jμ(x) ,which in this context is
more familiarly known as a current. We will assume that the current is conserved so that∂
μJμ=0.
1When I took a field theory course as a student with Sidney Coleman this was how he treated QED in order
to avoid discussing gauge invariance.
I.5. Coulomb and Newton | 33
Well, you know that the field theory of our vector meson is defined by the path integral
Z=/integraltext
DA eiS(A)≡eiW(J)with the action
S(A)=/integraldisplay
d4xL=/integraldisplay
d4x{1
2Aμ[(∂2+m2)gμν−∂μ∂ν]Aν+AμJμ} (1)
The second equality follows upon integrating by parts [compare (I.3.15)].
By now you have learned that we simply apply (I.3.16). We merely have to find the inverse
of the differential operator in the square bracket; in other words, we have to solve
[(∂2+m2)gμν−∂μ∂ν]Dνλ(x)=δμ
λδ(4)(x) (2)
As before [compare (I.3.17)] we go to momentum space by defining
Dνλ(x)=/integraldisplayd4k
(2π)4Dνλ(k)eikx
Plugging in, we find that [ −(k2−m2)gμν+kμkν]Dνλ(k)=δμ
λ, giving
Dνλ(k)=−gνλ+kνkλ/m2
k2−m2(3)
This is the photon, or more accurately the massive vector meson, propagator. Thus
W(J) =−1
2/integraldisplayd4k
(2π)4Jμ(k)∗−gμν+kμkν/m2
k2−m2+iεJν(k) (4)
Since current conservation ∂μJμ(x)=0 gets translated into momentum space as
kμJμ(k)=0, we can throw away the kμkνterm in the photon propagator. The effective
action simplifies to
W(J) =1
2/integraldisplayd4k
(2π)4Jμ(k)∗ 1
k2−m2+iεJμ(k) (5)
No further computation is needed to obtain a profound result. Just compare this result
to (I.4.2). The field theory has produced an extra sign. The potential energy between twolumps of charge density J
0(x) is positive. The electromagnetic force between like charges
is repulsive!
We can now safely let the photon mass mgo to zero thanks to current conservation.
[Note that we could not have done that in (3).] Indeed, referring to (I.4.7) we see that thepotential energy between like charges is
E=1
4πre−mr→1
4πr(6)
T o accommodate positive and negative charges we can simply write Jμ=Jμ
p−Jμ
n.W e
see that a lump with charge density J0
pis attracted to a lump with charge density J0
n.
Bypassing Maxwell
Having done electromagnetism in two minutes flat let me now do gravity. Let us move on
to the massive spin 2 meson field. In my treatment of the massive spin 1 meson field I
34 | I. Motivation and Foundation
took a short cut. Assuming that you are familiar with the Maxwell Lagrangian, I simply
added a mass term to it and took off. But I do not feel comfortable assuming that you areequally familiar with the corresponding Lagrangian for the massless spin 2 field (the so-called linearized Einstein Lagrangian, which I will discuss in a later chapter). So here I willfollow another strategy.
I invite you to think physically, and together we will arrive at the propagator for a massive
spin 2 field. First, we will warm up with the massive spin 1 case.
In fact, start with something even easier: the propagator D(k)=1/(k
2−m2)for a
massive spin 0 field. It tells us that the amplitude for the propagation of a spin 0 disturbanceblows up when the disturbance is almost a real particle. The residue of the pole is a propertyof the particle. The propagator for a spin 1 field D
νλcarries a pair of Lorentz indices and
in fact we know what it is from (3):
Dνλ(k)=−Gνλ
k2−m2(7)
where for later convenience we have defined
Gνλ(k)≡gνλ−kνkλ
m2(8)
Let us now understand the physics behind Gνλ. I expect you to remember the concept
of polarization from your course on electromagnetism. A massive spin 1 particle has threedegrees of polarization for the obvious reason that in its rest frame its spin vector can pointin three different directions. The three polarization vectors ε
(a)
μare simply the three unit
vectors pointing along the x,y, and zaxes, respectively (a =1, 2, 3 ):ε(1)
μ=(0, 1, 0, 0 ),
ε(2)
μ=(0, 0, 1, 0 ),ε(3)
μ=(0, 0, 0, 1 ). In the rest frame kμ=(m,0 ,0 ,0 )and so
kμε(a)
μ=0 (9)
Since this is a Lorentz invariant equation, it holds for a moving spin 1 particle as well.
Indeed, with a suitable normalization condition this fixes the three polarization vectorsε
(a)
μ(k) for a particle with momentum k.
The amplitude for a particle with momentum kand polarization ato be created at
the source is proportional to ε(a)
λ(k), and the amplitude for it to be absorbed at the sink
is proportional to ε(a)
ν(k). We multiply the amplitudes together to get the amplitude for
propagation from source to sink, and then sum over the three possible polarizations.Now we understand the residue of the pole in the spin 1 propagator D
νλ(k): It represents/summationtext
aε(a)
ν(k) ε(a)
λ(k). T o calculate this quantity, note that by Lorentz invariance it can only be
a linear combination of gνλandkνkλ. The condition kμε(a)
μ=0 fixes it to be proportional
togνλ−kνkλ/m2. We evaluate the left-hand side for kat rest with ν=λ=1, for instance,
and fix the overall and all-crucial sign to be −1. Thus
/summationdisplay
aε(a)
ν(k)ε(a)
λ(k)=−Gνλ(k)≡−/parenleftbigg
gνλ−kνkλ
m2/parenrightbigg
(10)
We have thus constructed the propagator Dνλ(k)for a massive spin 1 particle, bypassing
Maxwell (see appendix 1).
Onward to spin 2! We want to similarly bypass Einstein.
I.5. Coulomb and Newton | 35
Bypassing Einstein
A massive spin 2 particle has 5 (2 .2+1=5, remember?) degrees of polarization, char-
acterized by the five polarization tensors ε(a)
μν(a=1, 2, ... ,5)symmetric in the indices μ
andνsatisfying
kμε(a)
μν=0 (11)
and the tracelessness condition
gμνε(a)
μν=0 (12)
Let’s count as a check. A symmetric Lorentz tensor has 4 .5/2=10 components. The four
conditions in (11) and the single condition in (12) cut the number of components down to10−4−1=5, precisely the right number. (Just to throw some jargon around, remember
how to construct irreducible group representations? If not, read appendix B.) We fix thenormalization of ε
μνby setting the positive quantity/summationtext
aε(a)
12(k)ε(a)
12(k)=1.
So, in analogy with the spin 1 case we now determine/summationtext
aε(a)
μν(k)ε(a)
λσ(k). We have to
construct this object out of gμνandkμ, or equivalently Gμνandkμ. This quantity must
be a linear combination of terms such as GμνGλσ,Gμνkλkσ, and so forth. Using (11) and
(12) repeatedly (exercise I.5.1) you will easily find that
/summationdisplay
aε(a)
μν(k)ε(a)
λσ(k)=(GμλGνσ+GμσGνλ)−2
3GμνGλσ (13)
The overall sign and proportionality constant are determined by evaluating both sides for
kat rest (for μ=λ=1 and ν=σ=2, for instance).
Thus, we have determined the propagator for a massive spin 2 particle
Dμν,λσ(k)=(GμλGνσ+GμσGνλ)−2
3GμνGλσ
k2−m2(14)
Why we fall
We are now ready to understand one of the fundamental mysteries of the universe: Why
masses attract.
Recall from your courses on electromagnetism and special relativity that the energy or
mass density out of which mass is composed is part of a stress-energy tensor Tμν. For our
purposes, in fact, all you need to remember is that Tμνis a symmetric tensor and that
the component T00is the energy density. If you don’t remember, I will give you a physical
explanation in appendix 2.
T o couple to the stress-energy tensor, we need a tensor field ϕμνsymmetric in its two
indices. In other words, the Lagrangian of the world should contain a term like ϕμνTμν.
This is in fact how we know that the graviton, the particle responsible for gravity, has spin 2,just as we know that the photon, the particle responsible for electromagnetism and hence
36 | I. Motivation and Foundation
coupled to the current Jμ, has spin 1. In Einstein’s theory, which we will discuss in a later
chapter, ϕμνis of course part of the metric tensor.
Just as we pretended that the photon has a small mass to avoid having to discuss gauge
invariance, we will pretend that the graviton has a small mass to avoid having to discussgeneral coordinate invariance.
2Aha, we just found the propagator for a massive spin 2
particle. So let’s put it to work.
In precise analogy to (4)
W(J) =−1
2/integraldisplayd4k
(2π)4Jμ(k)∗−gμν+kμkν/m2
k2−m2+iεJν(k) (15)
describing the interaction between two electromagnetic currents, the interaction between
two lumps of stress energy is described by
W(T) =
−1
2/integraldisplayd4k
(2π)4Tμν(k)∗(GμλGνσ+GμσGνλ)−2
3GμνGλσ
k2−m2+iεTλσ(k)(16)
From the conservation of energy and momentum ∂μTμν(x)=0 and hence
kμTμν(k)=0, we can replace Gμνin (16) by gμν. (Here as is clear from the context gμνstill
denotes the flat spacetime metric of Minkowski, rather than the curved metric of Einstein.)
Now comes the punchline. Look at the interaction between two lumps of energy density
T00. We have from (16) that
W(T) =−1
2/integraldisplayd4k
(2π)4T00(k)∗1+1−2
3
k2−m2+iεT00(k) (17)
Comparing with (5) and using the well-known fact that (1+1−2
3)>0, we see that while
like charges repel, masses attract. T rumpets, please!
The universe
It is difficult to overstate the importance (not to speak of the beauty) of what we havelearned: The exchange of a spin 0 particle produces an attractive force, of a spin 1 particlea repulsive force, and of a spin 2 particle an attractive force, realized in the hadronic stronginteraction, the electromagnetic interaction, and the gravitational interaction, respectively.The universal attraction of gravity produces an instability that drives the formation ofstructure in the early universe.
3Denser regions become denser yet. The attractive nuclear
force mediated by the spin 0 particle eventually ignites the stars. Furthermore, the attractiveforce between protons and neutrons mediated by the spin 0 particle is able to overcomethe repulsive electric force between protons mediated by the spin 1 particle to form a
2For the moment, I ask you to ignore all subtleties and simply assume that in order to understand gravity it
is kosher to let m→0. I will give a precise discussion of Einstein’s theory of gravity in chapter VIII.1.
3A good place to read about gravitational instability and the formation of structure in the universe along the
line sketched here is in A. Zee, Einstein ’s Universe (formerly known as An Old Man ’s T oy ).
I.5. Coulomb and Newton | 37
variety of nuclei without which the world would certainly be rather boring. The repulsion
between likes and hence attraction between opposites generated by the spin 1 particle allowelectrically neutral atoms to form.
The world results from a subtle interplay among spin 0, 1, and 2.In this lightning tour of the universe, we did not mention the weak interaction. In fact,
the weak interaction plays a crucial role in keeping stars such as our sun burning at asteady rate.
Time differs from space by a sign
This weaving together of fields, particles, and forces to produce a universe rich withpossibilities is so beautiful that it is well worth pausing to examine the underlying physicssome more. The expression in (I.4.1) describes the effect of our disturbing the vacuum(or the mattress!) with the source J, calculated to second order. Thus some readers may
have recognized that the negative sign in (I.4.6) comes from the elementary quantummechanical result that in second order perturbation theory the lowest energy state alwayshas its energy pushed downward: for the ground state
all the energy denominators have the same sign.In essence, this “theorem” follows from the property of 2 by 2 matrices. Let us set the
ground state energy to 0 and crudely represent the entirety of the other states by a singlestate with energy w> 0. Then the Hamiltonian including the perturbation veffective to
second order is given by
H=/parenleftBiggwv
v 0/parenrightBigg
Since the determinant of H(and hence the product of the two eigenvalues) is manifestly
negative, the ground state energy is pushed below 0. [More explicitly, we calculate theeigenvalue εwith the characteristic equation 0 =ε(ε−w)−v
2≈−(wε+v2), and hence
ε≈−v2
w.] In different fields of physics, this phenomenon is variously known as level
repulsion or the seesaw mechanism (see chapter VII.7).
Disturbing the vacuum with a source lowers its energy. Thus it is easy to understand
that generically the exchange of a particles leads to an attractive force.
But then why does the exchange of a spin 1 particle produces a repulsion between like
objects? The secret lies in the profundity that space differs from time by a sign, namely,thatg
00=+ 1 while gii=− 1 fori=1, 2, 3. In (10), the left-hand side is manifestly positive
forν=λ=i. T aking kto be at rest we understand the minus sign in (10) and hence in (4).
Roughly speaking, for spin 2 exchange the sign occurs twice in (16).
Degrees of freedom
Now for a bit of cold water: Logically and mathematically the physics of a particle with mass
m/negationslash=0 could well be different from the physics with m=0. Indeed, we know from classical
38 | I. Motivation and Foundation
electromagnetism that an electromagnetic wave has 2 polarizations, that is, 2 degrees of
freedom. For a massive spin 1 particle we can go to its rest frame, where the rotation grouptells us that there are 2 .1+1=3 degrees of freedom. The crucial piece of physics is that we
can never bring the massless photon to its rest frame. Mathematically, the rotation groupSO( 3)degenerates into SO( 2), the group of 2-dimensional rotations around the direction
of the photon’s momentum.
We will see in chapter II.7 that the longitudinal degree of freedom of a massive spin 1
meson decouples as we take the mass to zero. The treatment given here for the interactionbetween charges (6) is correct. However, in the case of gravity, the
2
3in (17) is replaced by
1 in Einstein’s theory, as we will see chapter VIII.1. Fortunately, the sign of the interactiongiven in (17) does not change. Mute the trumpets a bit.
Appendix 1
Pretend that we never heard of the Maxwell Lagrangian. We want to construct a relativistic Lagrangian for a
massive spin 1 meson field. T ogether we will discover Maxwell. Spin 1 means that the field transforms as a vectorunder the 3-dimensional rotation group. The simplest Lorentz object that contains the 3-dimensional vector isobviously the 4-dimensional vector. Thus, we start with a vector field A
μ(x).
That the vector field carries mass mmeans that it satisfies the field equation
(∂2+m2)Aμ=0 (18)
A spin 1 particle has 3 degrees of freedom [remember, in fancy language, the representation jof the rotation
group has dimension (2j+1); here j=1.] On the other hand, the field Aμ(x) contains 4 components. Thus, we
must impose a constraint to cut down the number of degrees of freedom from 4 to 3. The only Lorentz covariantpossibility (linear in A
μ)is
∂μAμ=0 (19)
It may also be helpful to look at (18) and (19) in momentum space, where they read (k2−m2)Aμ(k)=0 and
kμAμ(k)=0. The first equation tells us that k2=m2and the second that if we go to the rest frame kμ=(m,/vector0)
thenA0vanishes, leaving us with 3 nonzero components Aiwithi=1, 2, 3.
The remarkable observation is that we can combine (18) and (19) into a single equation, namely
(gμν∂2−∂μ∂ν)Aν+m2Aμ=0 (20)
Verify that (20) contains both (18) and (19). Act with ∂μon (20). We obtain m2∂μAμ=0, which implies that
∂μAμ=0 . (At this step it is crucial that m/negationslash=0 and that we are not talking about the strictly massless photon.)
We have thus obtained (19 ); using (19) in (20) we recover (18).
We can now construct a Lagrangian by multiplying the left-hand side of (20) by +1
2Aμ(the1
2is conventional
but the plus sign is fixed by physics, namely the requirement of positive kinetic energy); thus
L=1
2Aμ[(∂2+m2)gμν−∂μ∂ν]Aν (21)
Integrating by parts, we recognize this as the massive version of the Maxwell Lagrangian. In the limit m→0w e
recover Maxwell.
A word about terminology: Some people insist on calling only Fμνa field and Aμa potential. Conforming to
common usage, we will not make this fine distinction. For us, any dynamical function of spacetime is a field.
I.5. Coulomb and Newton | 39
Appendix 2: Why does the graviton have spin 2?
First we have to understand why the photon has spin 1. Think physically. Consider a bunch of electrons at
rest inside a small box. An observer moving by sees the box Lorentz-Fitzgerald contracted and thus a highercharge density than the observer at rest relative to the box. Thus charge density J
0(x) transforms like the time
component of a 4-vector density Jμ(x). In other words, J/prime0=J0/√
1−v2. The photon couples to Jμ(x) and has
to be described by a 4-vector field Aμ(x) for the Lorentz indices to match.
What about energy density? The observer at rest relative to the box sees each electron contributing mto the
energy enclosed in the box. The moving observer, on the other hand, sees the electrons moving and thus each
having an energy m/√
1−v2. With the contracted volume and the enhanced energy, the energy density gets
enhanced by two factors of 1 /√
1−v2, that is, it transforms like the T00component of a 2-indexed tensor Tμν.
The graviton couples to Tμν(x) and has to be described by a 2-indexed tensor field ϕμν(x) for the Lorentz indices
to match.
Exercise
I.5.1 Write down the most general form for/summationtext
aε(a)
μν(k)ε(a)
λσ(k)using symmetry repeatedly. For example, it must
be invariant under the exchange {μν↔λσ}. You might end up with something like
AGμνGλσ+B(GμλGνσ+GμσGνλ)+C(Gμνkλkσ+kμkνGλσ)
+D(kμkλGνσ+kμkσGνλ+kνkσGμλ+kνkλGμσ)+Ekμkνkλkσ (22)
with various unknown A,... ,E. Apply kμ/summationtext
aε(a)
μν(k)ε(a)
λσ(k)=0 and find out what that implies for the
constants. Proceeding in this way, derive (13).
I.6 Inverse Square Law and the Floating 3-Brane
Why inverse square?
In your first encounter with physics, didn’t you wonder why an inverse square force law
and not, say, an inverse cube law? In chapter I.4 you learned the deep answer. When amassless particle is exchanged between two particles, the potential energy between thetwo particles goes as
V( r)∝/integraldisplay
d3kei/vectork./vectorx1
/vectork2∝1
r(1)
The spin of the exchanged particle controls the overall sign, but the 1 /rfollows just from
dimensional analysis, as I remarked earlier. Basically, V( r) is the Fourier transform of the
propagator. The /vectork2in the propagator comes from the (∂iϕ.∂iϕ)term in the action, where
ϕdenotes generically the field associated with the massless particle being exchanged, and
the(∂iϕ.∂iϕ)form is required by rotational invariance. It couldn’t be /vectorkor/vectork3in (1); /vectork2is
the simplest possibility. So you can say that in some sense ultimately the inverse squarelaw comes from rotational invariance!
Physically, the inverse square law goes back to Faraday’s flux picture. Consider a sphere
of radius rsurrounding a charge. The electric flux per unit area going through the sphere
varies as 1 /4πr
2. This geometric fact is reflected in the factor d3kin (1).
Brane world
Remarkably, with the tiny bit of quantum field theory I have exposed you to, I can already
take you to the frontier of current research, current as of the writing of this book. In stringtheory, our (3+1)-dimensional world could well be embedded in a larger universe, the
way a(2+1)-dimensional sheet of paper is embedded in our everyday (3+1)-dimensional
world. We are said to be living on a 3 brane.
So suppose there are nextra dimensions, with coordinates x
4,x5,... ,xn+3. Let the
characteristic scales associated with these extra coordinates be R. I can’t go into the
I.6. Inverse Square Law | 41
different detailed scenarios describing what Ris precisely. For some reason I can’t go into
either, we are stuck on the 3 brane. In contrast, the graviton is associated intrinsically withthe structure of spacetime and so roams throughout the (n+3+1)-dimensional universe.
All right, what is the gravitational force law between two particles? It is surely not your
grandfather’s gravitational force law: We Fourier transform
V( r)∝/integraldisplay
d3+nkei/vectork./vectorx1
/vectork2∝1
r1+n(2)
to obtain a 1 /r1+nlaw.
Doesn’t this immediately contradict observation?Well, no, because Newton’s law continues to hold for r/greatermuchR. In this regime, the extra
coordinates are effectively zero compared to the characteristic length scale rwe are inter-
ested in. The flux cannot spread far in the direction of the nextra coordinates. Think of
the flux being forced to spread in only the three spatial directions we know, just as theelectromagnetic field in a wave guide is forced to propagate down the tube. Effectively weare back in (3+1)-dimensional spacetime and V( r) reverts to a 1 /rdependence.
The new law of gravity (2) holds only in the opposite regime r/lessmuchR. Heuristically, when
Ris much larger than the separation between the two particles, the flux does not know
that the extra coordinates are finite in extent and thinks that it lives in an (n+3+1)-
dimensional universe.
Because of the weakness of gravity, Newton’s force law has not been tested to much
accuracy at laboratory distance scales, and so there is plenty of room for theorists tospeculate in: Rcould easily be much larger than the scale of elementary particles and yet
much smaller than the scale of everyday phenomena. Incredibly, the universe could have“large extra dimensions”! (The word “large” means large on the scale of particle physics.)
Planck mass
T o be quantitative, let us define the Planck mass MPlby writing Newton’s law more
rationally as V( r)=GNm1m2(1/r)=(m1m2/M2
Pl)(1/r). Numerically, MPl/similarequal1019Gev.
This enormous value obviously reflects the weakness of gravity.
In fundamental units in which /planckover2piandcare set to unity, gravity defines an intrinsic mass
or energy scale much higher than any scale we have yet explored experimentally. Indeed,one of the fundamental mysteries of contemporary particle physics is why this mass scaleis so high compared to anything else we know of. I will come back to this so-called hierarchyproblem in due time. For the moment, let us ask if this new picture of gravity, new in thewaning moments of the last century, can alleviate the hierarchy problem by lowering theintrinsic mass scale of gravity.
Denote the mass scale (the “true scale” of gravity) characteristic of gravity in the (n+3
+1)-dimensional universe by M
TGso that the gravitational potential between two objects
of masses m1andm2separated by a distance r/lessmuchRis given by
V( r)=m1m2
[MTG]2+n1
r1+n
42 | I. Motivation and Foundation
Note that the dependence on MTGfollows from dimensional analysis: two powers to cancel
m1m2andnpowers to match the nextra powers of 1 /r.F o rr/greatermuchR, as we have argued, the
geometric spread of the gravitational flux is cut off by Rso that the potential becomes
V( r)=m1m2
[MTG]2+n1
Rn1
r
Comparing with the observed law V( r)=(m 1m2/M2
Pl)(1/r) we obtain
M2
TG=M2
Pl
[MTGR]n(3)
IfMTGRcould be made large enough, we have the intriguing possibility that the funda-
mental scale of gravity MTGmay be much lower than what we have always thought.
Accelerators (such as the large Hadron Collider) could offer an exciting verification of
this intriguing possibility. If the true scale of gravity MTGlies in an energy range accessible
to the accelerator, there may be copious production of gravitons escaping into the higherdimensional universe. Experimentalists would see a massive amount of missing energy.
Exercise
I.6.1 Putting in the numbers, show that the case n=1 is already ruled out.
I.7 Feynman Diagrams
Feynman brought quantum field theory to the masses.
—J. Schwinger
Anharmonicity in field theory
The free field theory we studied in the last few chapters was easy to solve because the defin-ing path integral (I.3.14) is Gaussian, so we could simply apply (I.2.15). (This correspondsto solving the harmonic oscillator in quantum mechanics.) As I noted in chapter I.3, withinthe harmonic approximation the vibrational modes on the mattress can be linearly super-posed and thus they simply pass through each other. The particles represented by wavepackets constructed out of these modes do not interact:
1hence the term free field theory.
T o have the modes scatter off each other we have to include anharmonic terms in the La-grangian so that the equation of motion is no longer linear. For the sake of simplicity letus add only one anharmonic term −
λ
4!ϕ4to our free field theory and, recalling (I.3.11), try
to evaluate
Z(J)=/integraldisplay
Dϕ ei/integraltext
d4x{1
2[(∂ϕ)2−m2ϕ2]−λ
4!ϕ4+Jϕ}(1)
(We suppress the dependence of Zonλ.)
Doing quantum field theory is no sweat, you say, it just amounts to doing the functional
integral (1). But the integral is not easy! If you could do it, it would be big news.
Feynman diagrams made easy
As an undergraduate, I heard of these mysterious little pictures called Feynman diagramsand really wanted to learn about them. I am sure that you too have wondered about those
1A potential source of confusion: Thanks to the propagation of ϕ, the sources coupled to ϕinteract, as was
seen in chapter I.4, but the particles associated with ϕdo not interact with each other. This is like saying that
charged particles coupled to the photon interact, but (to leading approximation) photons do not interact with
each other.
44 | I. Motivation and Foundation
funny diagrams. Well, I want to show you that Feynman diagrams are not such a big deal:
Indeed we have already drawn little spacetime pictures in chapters I.3 and I.4 showinghow particles can appear, propagate, and disappear.
Feynman diagrams have long posed somewhat of an obstacle for first-time learners
of quantum field theory. T o derive Feynman diagrams, traditional texts typically adopt thecanonical formalism (which I will introduce in the next chapter) instead of the path integralformalism used here. As we will see, in the canonical formalism fields appear as quantumoperators. T o derive Feynman diagrams, we would have to solve the equation of motionof the field operators perturbatively in λ. A formidable amount of machinery has to be
developed.
In the opinion of those who prefer the path integral, the path integral formalism
derivation is considerably simpler (naturally!). Nevertheless, the derivation can still getrather involved and the student could easily lose sight of the forest for the trees. There isno getting around the fact that you would have to put in some effort.
I will try to make it as easy as possible for you. I have hit upon the great pedagogical
device of letting you discover the Feynman diagrams for yourself. My strategy is to let youtackle two problems of increasing difficulty, what I call the baby problem and the childproblem. By the time you get through these, the problem of evaluating (1) will seem muchmore tractable.
A baby problem
The baby problem is to evaluate the ordinary integral
Z(J)=/integraldisplay+∞
−∞dqe−1
2m2q2−λ
4!q4+Jq(2)
evidently a much simpler version of (1).
First, a trivial point: we can always scale q→q/m so that Z=m−1F(λ
m4,J
m), but we
won’t.
Forλ=0 this is just one of the Gaussian integrals done in the appendix of chapter I.2.
Well, you say, it is easy enough to calculate Z(J) as a series in λ: expand
Z(J)=/integraldisplay+∞
−∞dqe−1
2m2q2+Jq/bracketleftbigg
1−λ
4!q4+1
2(λ
4!)2q8+.../bracketrightbigg
and integrate term by term. You probably even know one of several tricks for computing/integraltext+∞
−∞dqe−1
2m2q2+Jqq4n: you write it as (d
dJ)4n/integraltext+∞
−∞dqe−1
2m2q2+Jqand refer to (I.2.19). So
Z(J)=(1−λ
4!(d
dJ)4+1
2(λ
4!)2(d
dJ)8+...)/integraldisplay+∞
−∞dqe−1
2m2q2+Jq(3)
=e−λ
4!(d
dJ)4/integraldisplay+∞
−∞dqe−1
2m2q2+Jq=(2π
m2)1
2e−λ
4!(d
dJ)4e1
2m2J2
(4)
(There are other tricks, such as differentiating/integraltext+∞
−∞dqe−1
2m2q2+Jqwith respect to m2
repeatedly, but I want to discuss a trick that will also work for field theory.) By expanding
I.7. Feynman Diagrams | 45
J J JJ JJ
J J JJ JJλλ λ
(a) (b) (c)
( )4λJ4 1
m2
Figure I.7.1
the two exponentials we can obtain any term in a double series expansion of Z(J) inλand
J. [We will often suppress the overall factor (2π/m2)1
2=Z(J=0,λ=0)≡Z(0, 0)since it
will be common to all terms. When we want to be precise, we will define ˜Z=Z(J)/Z(0, 0 ).]
For example, suppose we want the term of order λandJ4in˜Z. We extract the order
J8term in eJ2/2m2, namely, [1 /4!(2m2)4]J8, replace e−(λ/4! )(d/dJ)4by−(λ/4! )(d/dJ)4,
and differentiate to get [8! (−λ)/(4! )3(2m2)4]J4. Another example: the term of order λ2
andJ4is [12! (−λ)2/(4!)36!2(2m2)6]J4. A third example: the term of order λ2andJ6is
1
2(λ/4!)2(d/dJ)8[1/7!(2m2)7]J14=[14!(−λ)2/(4!)26!7!2(2m2)7]J6. Finally, the term of order
λandJ0is [1/2(2m2)2](−λ) .
You can do this as well as I can! Do a few more and you will soon see a pattern. In
fact, you will eventually realize that you can associate diagrams with each term and codifysome rules. Our four examples are associated with the diagrams in figures I.7.1–I.7.4.You can see, for a reason you will soon understand, that each term can be associated withseveral diagrams. I leave you to work out the rules carefully to get the numerical factorsright (but trust me, the “future of democracy” is not going to depend on them). The rulesgo something like this: (1) diagrams are made of lines and vertices at which four linesmeet; (2) for each vertex assign a factor of (−λ); (3) for each line assign 1 /m
2; and (4) for
each external end assign J(e.g., figure I.7.3 has seven lines, two vertices, and six ends,
giving ∼[(−λ)2/(m2)7]J6.) (Did you notice that twice the number of lines is equal to four
times the number of vertices plus the number of ends? We will meet relations like that inchapter III.2.)
In addition to the two diagrams shown in figure I.7.3, there are ten diagrams obtained
by adding an unconnected straight line to each of the ten diagrams in figure I.7.2. (Do youunderstand why?)
For obvious reasons, some diagrams (e.g., figure I.7.1a, I.7.3a) are known as tree
2
diagrams and others (e.g., Figs. I.7.1b and I.7.2a) as loop diagrams.
Do as many examples as you need until you feel thoroughly familiar with what is going
on, because we are going to do exactly the same thing in quantum field theory. It willlook much messier, but only superficially. Be sure you understand how to use diagrams to
2The Chinese character for tree (A. Zee, Swallowing Clouds ) is shown in fig. I.7.5. I leave it to you to figure
out why this diagram does not appear in our Z(J) .
J J
J J/H9261
(a) (b) (d)
(e)
(i) (j)
( )6/H92612J4(f ) (g) (h)(c)/H9261
1
m2
Figure I.7.2
J J
J J/H9261
(a) (b)
( )7/H92612J6 1
m2J
/H9261
J
Figure I.7.3
Figure I.7.4
I.7. Feynman Diagrams | 47
Figure I.7.5
represent the double series expansion of ˜Z(J) before reading on. Please. In my experience
teaching, students who have not thoroughly understood the expansion of ˜Z(J) have no
hope of understanding what we are going to do in the field theory context.
Wick contraction
It is more obvious than obvious that we can expand Z(J) in powers of J, if we please,
instead of in powers of λ. As you will see, particle physicists like to classify in power of J.
In our baby problem, we can write
Z(J)=∞/summationdisplay
s=01
s!Js/integraldisplay+∞
−∞dqe−1
2m2q2−(λ/4! )q4qs≡Z(0, 0)∞/summationdisplay
s=01
s!JsG(s)(5)
The coefficient G(s), whose analogs are known as “Green’s functions” in field theory, can
be evaluated as a series in λwith each term determined by Wick contraction (I.2.10). For
instance, the O(λ) term in G(4)is
−λ
4!Z(0, 0)/integraldisplay+∞
−∞dqe−1
2m2q2q8=−7!!
4!1
m8
which of course better be equal3to what we obtained above for the λJ4term in ˜Z. Thus,
there are two ways of computing Z: you expand in λfirst or you expand in Jfirst.
Connected versus disconnected
You will have noticed that some Feynman diagrams are connected and others are not.
Thus, figure I.7.3a is connected while 3b is not. I presaged this at the end of chapter I.4and in figures I.4.2 and I.4.3. Write
Z(J ,λ)=Z(J=0,λ)eW(J ,λ)=Z(J=0,λ)∞/summationdisplay
N=01
N![W(J ,λ)]N(6)
By definition, Z(J=0,λ)consists of those diagrams with no external source J, such as
the one in figure I.7.4. The statement is that Wis a sum of connected diagrams while
3As a check on the laws of arithmetic we verify that indeed 7!! /(4!)2=8!/(4!)324.
48 | I. Motivation and Foundation
Zcontains connected as well as disconnected diagrams. Thus, figure I.7.3b consists of
two disconnected pieces and comes from the term (1/2!)[W(J ,λ)]2in (6), the 2! taking
into account that it does not matter which of the two pieces you put “on the left or on theright.” Similarly, figure I.7.2i comes from (1/3!)[W(J ,λ)]
3. Thus, it is Wthat we want to
calculate, not Z. If you’ve had a good course on statistical mechanics, you will recognize
that this business of connected graphs versus disconnected graphs is just what underliesthe relation between free energy and the partition function.
Propagation: from here to there
All these features of the baby problem are structurally the same as the correspondingfeatures of field theory and we can take over the discussion almost immediately. But beforewe graduate to field theory, let us consider what I call a child problem, the evaluation of amultiple integral instead of a single integral:
Z(J)=/integraldisplay+∞
−∞/integraldisplay+∞
−∞.../integraldisplay+∞
−∞dq1dq2...dqNe−1
2q.A.q−(λ/4! )q4+J.q(7)
withq4≡/summationtext
iq4
i. Generalizing the steps leading to (3) we obtain
Z(J)=/bracketleftbigg(2π)N
det[A]/bracketrightbigg1
2
e−(λ/4! )/summationtext
i(∂/∂J i)4e1
2J.A−1.J(8)
Alternatively, just as in (5) we can expand in powers of J
Z(J)=∞/summationdisplay
s=0N/summationdisplay
i1=1...N/summationdisplay
is=11
s!Ji1...Jis/integraldisplay+∞
−∞/parenleftBigg/productdisplay
ldql/parenrightBigg
e−1
2q.A.q−(λ/4! )q4qi1...qis
=Z(0, 0)∞/summationdisplay
s=0N/summationdisplay
i1=1...N/summationdisplay
is=11
s!Ji1...JisG(s)
i1...is(9)
which again we can expand in powers of λand evaluate by Wick contracting.
The one feature the child problem has that the baby problem doesn’t is propagation
“from here to there”. Recall the discussion of the propagator in chapter I.3. Just as in(I.3.16) we can think of the index ias labeling the sites on a lattice. Indeed, in (I.3.16) we
had in effect evaluated the “2-point Green’s function” G
(2)
ijto zeroth order in λ(differentiate
(I.3.16) with respect to Jtwice):
G(2)
ij(λ=0)=/bracketleftBigg/integraldisplay+∞
−∞/parenleftBigg/productdisplay
ldql/parenrightBigg
e−1
2q.A.qqiqj/bracketrightBigg
/Z(0, 0 )=(A−1)ij
(see also the appendix to chapter I.2). The matrix element ( A−1)ijdescribes propagation
fromitoj. In the baby problem, each term in the expansion of Z(J) can be associated
with several diagrams but that is no longer true with propagation.
I.7. Feynman Diagrams | 49
Let us now evaluate the “4-point Green’s function” G(4)
ijklto order λ:
G(4)
ijkl=/integraldisplay+∞
−∞/parenleftBigg/productdisplay
mdqm/parenrightBigg
e−1
2q.A.qqiqjqkql/bracketleftBigg
1−λ
4!/summationdisplay
nq4
n+O(λ2)/bracketrightBigg
/Z(0, 0 )
=(A−1)ij(A−1)kl+(A−1)ik(A−1)jl+(A−1)il(A−1)jk
−λ/summationdisplay
n(A−1)in(A−1)jn(A−1)kn(A−1)ln+...+O(λ2) (10)
The first three terms describe one excitation propagating from itojand another propa-
gating from ktol, plus the two possible permutations on this “history.” The order λterm
tells us that four excitations, propagating from iton, from jton, from kton, and from l
ton, meet at nand interact with an amplitude proportional to λ, where nis anywhere on
the lattice or mattress. By the way, you also see why it is convenient to define the interac-tion(λ/4!)ϕ
4with a 1 /4! :qihas a choice of four qn’s to contract with, qjhas three qn’s to
contract with, and so on, producing a factor of 4! to cancel the (1/4!).
I intentionally did not display in (10) the O(λ) terms produced by Wick contracting some
of the qn’s with each other. There are two types: (I) Contracting a pair of qn’s produces
something like λ(A−1)ij(A−1)kn(A−1)ln(A−1)nnand (II) contracting the qn’s with each
other produces the first three terms in (10) multiplied by (A−1)nn(A−1)nn. We see that
(I) and (II) correspond to diagrams b and c in figure I.7.1, respectively. Evidently, the twoexcitations do not interact with each other. I will come back to (II) later in this chapter.
Perturbative field theory
You should now be ready for field theory!
Indeed, the functional integral in (1) (which I repeat here)
Z(J)=/integraldisplay
Dϕ ei/integraltext
d4x{1
2[(∂ϕ)2−m2ϕ2]−(λ/4! )ϕ4+Jϕ}(11)
has the same form as the ordinary integral in (2) and the multiple integral in (7). There
is one minor difference: there is no iin (2) and (7), but as I noted in chapter I.2 we can
Wick rotate (11) and get rid of the i, but we won’t. The significant difference is that Jand
ϕin (11) are functions of a continuous variable x, while Jandqin (2) are not functions of
anything and in (7) are functions of a discrete variable. Aside from that, everything goesthrough the same way.
As in (3) and (8) we have
Z(J)=e−(i/4! )λ/integraltext
d4w[δ/iδJ(w) ]4/integraldisplay
Dϕei/integraltext
d4x{1
2[(∂ϕ)2−m2ϕ2]+Jϕ}
=Z(0, 0)e−(i/4! )λ/integraltext
d4w[δ/iδJ(w) ]4e−(i/2)/integraltext/integraltext
d4xd4yJ(x)D(x −y)J(y)(12)
The structural similarity is total.
The role of 1 /m2in (3) and of A−1(8) is now played by the propagator
D(x−y)=/integraldisplayd4k
(2π)4eik.(x−y)
k2−m2+iε
50 | I. Motivation and Foundation
Incidentally, if you go back to chapter I.3 you will see that if we were in d-dimensional
spacetime, D(x−y)would be given by the same expression with d4k/(2π)4replaced by
ddk/(2π)d. The ordinary integral (2) is like a field theory in 0-dimensional spacetime: if
we set d=0, there is no propagating around and D(x−y)collapses to −1/m2. You see
that it all makes sense.
We also know that J(x) corresponds to sources and sinks. Thus, if we expand Z(J)
as a series in J, the powers of Jwould indicate the number of particles involved in the
process. (Note that in this nomenclature the scattering process ϕ+ϕ→ϕ+ϕcounts as a
4-particle process: we count the total number of incoming and outgoing particles.) Thus,in particle physics it often makes sense to specify the power of J. Exactly as in the baby
and child problems, we can expand in Jfirst:
Z(J)=Z(0, 0)∞/summationdisplay
s=0is
s!/integraldisplay
dx1...dxsJ(x 1)...J(xs)G(s)(x1,... ,xs)
=∞/summationdisplay
s=0is
s!/integraldisplay
dx1...dxsJ(x 1)...J(xs)/integraldisplay
Dϕ ei/integraltext
d4x{1
2[(∂ϕ)2−m2ϕ2]−(λ/4! )ϕ4}
ϕ(x1)...ϕ(xs) (13)
In particular, we have the 2-point Green’s function
G(x 1,x2)≡1
Z(0, 0)/integraldisplay
Dϕ ei/integraltext
d4x{1
2[(∂ϕ)2−m2ϕ2]−(λ/4! )ϕ4}ϕ(x1)ϕ(x 2) (14)
the 4-point Green’s function,
G(x1,x2,x3,x4)≡1
Z(0, 0)/integraldisplay
Dϕ ei/integraltext
d4x{1
2[(∂ϕ)2−m2ϕ2]−(λ/4! )ϕ4}ϕ(x 1)ϕ(x 2)ϕ(x 3)ϕ(x 4) (15)
and so on. [Sometimes Z(J) is called the generating functional as it generates the Green’s
functions.] Obviously, by translation invariance, G(x 1,x2)does not depend on x1andx2
separately, but only on x1−x2. Similarly, G(x1,x2,x3,x4)only depends on x1−x4,x2−x4,
andx3−x4.F o rλ=0,G(x1,x2)reduces to iD(x1−x2), the propagator introduced in
chapter I.3. While D(x1−x2)describes the propagation of a particle between x1andx2in
the absence of interaction, G(x 1−x2)describes the propagation of a particle between x1
andx2in the presence of interaction. If you understood our discussion of G(4)
ijkl, you would
know that G(x 1,x2,x3,x4)describes the scattering of particles.
In some sense, there are two ways of doing field theory, what I might call the Schwinger
way (12) or the Wick way (13).
Thus, to summarize, Feynman diagrams are just an extremely convenient way of rep-
resenting the terms in a double series expansion of Z(J) inλandJ.
As I said in the preface, I have no intention of turning you into a whiz at calculating
Feynman diagrams. In any case, that can only come with practice. Instead, I tried to giveyou as clear an account as I can muster of the concept behind this marvellous invention of
I.7. Feynman Diagrams | 51
3 4
1 2
xt
Figure I.7.6
Feynman’s, which as Schwinger noted rather bitterly, enables almost anybody to become
a field theorist. For the moment, don’t worry too much about factors of 4! and 2!
Collision between particles
As I already mentioned, I described in chapter I.4 the strategy of setting up sources andsinks to watch the propagation of a particle (which I will call a meson) associated withthe field ϕ. Let us now set up two sources and two sinks to watch two mesons scatter off
each other. The setup is shown in figure I.7.6. The sources localized in regions 1 and 2both produce a meson, and the two mesons eventually disappear into the sinks localizedin regions 3 and 4. It clearly suffices to find in Za term containing J(x
1)J(x 2)J(x3)J(x4).
But this is just G(x1,x2,x3,x4).
Let us be content with first order in λ. Going the Wick way we have to evaluate
1
Z(0, 0)/parenleftbigg
−iλ
4!/parenrightbigg/integraldisplay
d4w/integraldisplay
Dϕ ei/integraltext
d4x{1
2[(∂ϕ)2−m2ϕ2]}
ϕ(x 1)ϕ(x2)ϕ(x3)ϕ(x4)ϕ(w)4(16)
Just as in (10) we Wick contract and obtain
(−iλ)/integraldisplay
d4wD(x 1−w)D(x 2−w)D(x 3−w)D(x 4−w) (17)
As a check, let us also derive this the Schwinger way. Replace e−(i/4! )λ/integraltext
d4w(δ/δJ(w))4by
−(i/4! )λ/integraltext
d4w(δ/δJ(w))4ande−(i/2)/integraltext/integraltext
d4xd4yJ(x)D(x −y)J(y)by
i4
4!24/bracketleftbigg/integraldisplay/integraldisplay
d4xd4yJ(x)D(x −y)J(y)/bracketrightbigg4
52 | I. Motivation and Foundation
x3 x4
x1 x2k3 k4
k1 k2
(a) (b)w
Figure I.7.7
To save writing, it would be sagacious to introduce the abbreviations JaforJ(xa),/integraltext
afor/integraltext
d4xa, andDabforD(xa−xb). Dropping overall numerical factors, which I invite you to
fill in, we obtain
∼iλ/integraldisplay
w(δ
δJw)4/integraldisplay/integraldisplay/integraldisplay/integraldisplay/integraldisplay/integraldisplay/integraldisplay/integraldisplay
DaeDbfDcgDdhJaJbJcJdJeJfJgJh (18)
The four (δ/δJw)’s hit the eight J’s in all possible combinations producing many terms,
which again I invite you to write out. Two of the three terms are disconnected. Theconnected term is
∼iλ/integraldisplay
w/integraldisplay/integraldisplay/integraldisplay/integraldisplay
DawDbwDcwDdwJaJbJcJd (19)
Evidently, this comes from the four (δ/δJ w)’s hitting Je,Jf,Jg, and Jh, thus setting
xe,xf,xg, andxhtow. Compare (19) with [8! (−λ)/(4! )3(2m2)4]J4in the baby problem.
Recall that we embarked on this calculation in order to produce two mesons by sources
localized in regions 1 and 2, watch them scatter off each other, and then get rid of themwith sinks localized in regions 3 and 4. In other words, we set the source function J(x)
equal to a set of delta functions peaked at x
1,x2,x3, and x4. This can be immediately
read off from (19): the scattering amplitude is just −iλ/integraltext
wD1wD2wD3wD4w, exactly as
in (17).
The result is very easy to understand (see figure I.7.7a). Two mesons propagate from
their birthplaces at x1andx2to the spacetime point w, with amplitude D(x1−w)D(x2−
w), scatter with amplitude −iλ , and then propagate from wto their graves at x3andx4,
with amplitude D(w−x3)D(w −x4)[note that D(x)=D(−x) ]. The integration over w
just says that the interaction point wcould have been anywhere. Everything is as in the
child problem.
It is really pretty simple once you get it. Still confused? It might help if you think of (12)
as some kind of machine e−(i/4! )λ/integraltext
d4w[δ/iδJ(w) ]4operating on
Z(J ,λ=0)=e−(i/2)/integraltext/integraltext
d4xd4yJ(x)D(x −y)J(y)
When expanded out, Z(J ,λ=0)is a bunch of J’s thrown here and there in spacetime,
with pairs of J’s connected by D’s. Think of a bunch of strings with the string ends
I.7. Feynman Diagrams | 53
x3 x4
x1 x2
(a)y4 y3
y1 y2
x1 x2x3 x4
w
x3 x4
x1 x2y4 y3
y1 y2
x1 x2x3 x4
w1u2
z2u1
z1
(b)w2
Figure I.7.8
corresponding to the J’s. What does the “machine” do? The machine is a sum of terms,
for example, the term
∼λ2/integraldisplay
d4w1/integraldisplay
d4w2/bracketleftbiggδ
δJ(w 1)/bracketrightbigg4/bracketleftbiggδ
δJ(w 2)/bracketrightbigg4
When this term operates on a term in Z(J ,λ=0)it grabs four string ends and glues them
together at the point w2; then it grabs another 4 string ends and glues them together at
the point w1. The locations w1andw2are then integrated over. Two examples are shown
in figure I.7.8. It is a game you can play with a child! This childish game of gluing four
string ends together generates all the Feynman diagrams of our scalar field theory.
Do it once and for all
Now Feynman comes along and says that it is ridiculous to go through this long-winded
yakkety-yak every time. Just figure out the rules once and for all.
54 | I. Motivation and Foundation
For example, for the diagram in figure I.7.7a associate the factor −iλ with the scattering,
the factor D(x1−w)with the propagation from x1tow, and so forth—conceptually exactly
the same as in our baby problem. See, you could have invented Feynman diagrams. (Well,not quite. Maybe not, maybe yes.)
Just as in going from (I.4.1) to (I.4.2), it is easier to pass to momentum space. Indeed,
that is how experiments are actually done. A meson with momentum k
1and a meson with
momentum k2collide and scatter, emerging with momenta k3andk4(see figure I.7.7b).
Each spacetime propagator gives us
D(xa−w)=/integraldisplayd4ka
(2π)4e±ika(xa−w)
k2
a−m2+iε
Note that we have the freedom of associating with the dummy integration variable either
a plus or a minus sign in the exponential. Thus in integrating over win (17) we obtain
/integraldisplay
d4we−i(k 1+k 2−k 3−k 4)w=(2π)4δ(4)(k1+k2−k3−k4).
That the interaction could occur anywhere in spacetime translates into momentum con-
servation k1+k2=k3+k4. (We put in the appropriate minus signs in two of the D’s so
that we can think of k3andk4as outgoing momenta.)
So there are Feynman diagrams in real spacetime and in momentum space. Spacetime
Feynman diagrams are literally pictures of what happened. (A trivial remark: the orien-tation of Feynman diagrams is a matter of idiosyncratic choice. Some people draw themwith time running vertically, others with time running horizontally. We follow Feynmanin this text.)
The rules
We have thus derived the celebrated momentum space Feynman rules for our scalar fieldtheory:
1. Draw a Feynman diagram of the process (fig. I.7.7b for the example we just discussed).
2. Label each line with a momentum kand associate with it the propagator i/(k2−m2+iε).
3. Associate with each interaction vertex the coupling −iλ and(2π)4δ(4)(/summationtext
iki−/summationtext
jkj),
forcing the sum of momenta/summationtext
ikicoming into the vertex to be equal to the sum of momenta
/summationtext
jkjgoing out of the vertex.
4. Momenta associated with internal lines are to be integrated over with the measured4k
(2π)4.
Incidentally, this corresponds to the summation over intermediate states in ordinary per-turbation theory.
5. Finally, there is a rule about some really pesky symmetry factors. They are the analogs of
those numerical factors in our baby problem. As a result, some diagrams are to be multipliedby a symmetry factor such as
1
2. These originate from various combinatorial factors counting
the different ways in which the (δ/δJ )’s can hit the J’s in expressions such as (18). I will let
you discover a symmetry factor in exercise I.7.2.
We will illustrate by examples what these rules (and the concept of internal lines) mean.
I.7. Feynman Diagrams | 55
Our first example is just the diagram (fig. I.7.7b) that we have calculated. Applying the
rules we obtain
−iλ( 2π)4δ(4)(k1+k2−k3−k4)4/productdisplay
a=1/parenleftBigg
i
k2
a−m2+iε/parenrightBigg
You would agree that it is silly to drag the factor/producttext4
a=1/parenleftbigg
i
k2a−m2+iε/parenrightbigg
around, since it would
be common to all diagrams in which two mesons scatter into two mesons. So we append
to the Feynman rules an additional rule that we do not associate a propagator with externallines. (This is known in the trade as “amputating the external legs”.)
In an actual scattering experiment, the external particles are of course real and on shell,
that is, their momenta satisfy k
2
a−m2=0. Thus we better not keep the propagators of the
external lines around. Arithmetically, this amounts to multiplying the Green’s functions[and what we have calculated thus far are indeed Green’s functions, see (16)] by the factor/Pi1
a(−i)(k2
a−m2)and then set k2
a=m2(“putting the particles on shell” in conversation). At
this point, this procedure sounds like formal overkill. We will come back to the rationalebehind it at the end of the next chapter.
Also, since there is always an overall factor for momentum conservation we should not
drag the overall delta function around either. Thus, we have two more rules:
6. Do not associate a propagator with external lines.
7. A delta function for overall momentum conservation is understood.
Applying these rules we obtain an amplitude which we will denote by M. For example,
for the diagram in figure I.7.7b M=−iλ.
The birth of particles
As explained in chapter I.1 one of the motivations for constructing quantum field theory
was to describe the production of particles. We are now ready to describe how two col-liding mesons can produce four mesons. The Feynman diagram in figure I.7.9 (comparefig. I.7.3a) can occur in order λ
2in perturbation theory. Amputating the external legs, we
drop the factor/producttext6
a=1[i/k2
a−m2+iε] associated with the six external lines, keeping only
the propagator associated with the one internal line. For each vertex we put in a factor of(−iλ) and a momentum conservation delta function (rule 3). Then we integrate over the
internal momentum q(rule 4) to obtain
(−iλ)2/integraldisplayd4q
(2π)4i
q2−m2+iε(2π)4δ(4)(k1+k2−k3−q)(2π)4δ(4)[q−(k4+k5+k6)] (20)
The integral over qis a cinch, giving
(−iλ)2 i
(k4+k5+k6)2−m2+iε(2π)4δ(4)[k1+k2−(k3+k4+k5+k6)] (21)
56 | I. Motivation and Foundation
k3k4
k1 k2k5
k6
q
Figure I.7.9
We have already agreed (rule 7) not to drag the overall delta function around. This exam-
ple teaches us that we didn’t have to write down the delta functions and then annihilate (allbut one of) them by integrating. In figure I.7.9 we could have simply imposed momentumconservation from the beginning and labeled the internal line as k
4+k5+k6instead of q.
With some practice you could just write down the amplitude
M=(−iλ)2 i
(k4+k5+k6)2−m2+iε(22)
directly: just remember, a coupling (−iλ) for each vertex and a propagator for each internal
(but not external) line. Pretty easy once you get the hang of it. As Schwinger said, the massescould do it.
The cost of not being real
The physics involved is also quite clear: The internal line is associated with a virtual particle
whose relativistic 4-momentum k4+k5+k6squared is not necessarily equal to m2,a si t
would have to be if the particle were real. The farther the momentum of the virtual particle
is from the mass shell the smaller the amplitude. You are penalized for not being real.
According to the quantum rules for dealing with identical particles, to obtain the full
amplitude we have to symmetrize among the four final momenta. One way of saying it isto note that the line labeled k
3in figure (I.7.9) could have been labeled k4,k5,o rk6, and
we have to add all four contributions.
To make sure that you understand the Feynman rules I insist that you go through the
path integral calculation to obtain (21) starting with (12) and (13).
I am repeating myself but I think it is worth emphasizing again that there is nothing
particularly magical about Feynman diagrams.
I.7. Feynman Diagrams | 57
k1 k2k3 k4
k k1 + k2 − k
Figure I.7.10
Loops and a first look at divergence
Just as in our baby problem, we have tree diagrams and loop diagrams. So far we have only
looked at tree diagrams. Our next example is the loop diagram in figure I.7.10 (comparefig. I.7.2a.) Applying the Feynman rules, we obtain
1
2(−iλ)2/integraldisplayd4k
(2π)4i
k2−m2+iεi
(k1+k2−k)2−m2+iε(23)
As above, the physics embodied in (23) is clear: As kranges over all possible values, the
integrand is large only if one or the other or both of the virtual particles associated withthe two internal lines are close to being real. Once again, there is a penalty for not beingreal (see exercise I.7.4).
For large kthe integrand goes as 1 /k
4. The integral is infinite! It diverges as/integraltext
d4k(1/k4).
We will come back to this apparent disaster in chapter III.1.
With some practice, you will be able to write down the amplitude by inspection. As
another example, consider the three-loop diagram in figure I.7.11 contributing in O(λ4)to
meson-meson scattering. First, for each loop pick an internal line and label the momentum
it carries, p,q, andrin our example. There is considerable freedom of choice in labeling—
your choice may well not agree with mine, but of course the physics should not depend
on it. The momenta carried by the other internal lines are then fixed by momentumconservation, as indicated in the figure. Write down a coupling for each vertex, and apropagator for each internal line, and integrate over the internal momenta p,q, andr.
58 | I. Motivation and Foundation
k1k2k3 k4
r
pqp − q − rk1 + k2 − r
k1 + k2 − p
Figure I.7.11
Thus, without worrying about symmetry factors, we have the amplitude
(−iλ)4/integraldisplayd4p
(2π)4d4q
(2π)4d4r
(2π)4i
p2−m2+iεi
(k1+k2−p)2−m2+iε
i
q2−m2+iεi
(p−q−r)2−m2+iεi
r2−m2+iεi
(k1+k2−r)2−m2+iε(24)
Again, this triple integral also diverges: It goes as/integraltext
d12P(1/P12).
An assurance
When I teach quantum field theory, at this point in the course some students get un-
accountably very anxious about Feynman diagrams. I would like to assure the reader that
the whole business is really quite simple. Feynman diagrams can be thought of simply aspictures in spacetime of the antics of particles, coming together, colliding and producingother particles, and so on. One student was puzzled that the particles do not move in
I.7. Feynman Diagrams | 59
straight lines. Remember that a quantum particle propagates like a wave; D(x−y)gives
us the amplitude for the particle to propagate from xtoy. Evidently, it is more convenient to
think of particles in momentum space: Fourier told us so. We will see many more examplesof Feynman diagrams, and you will soon be well acquainted with them. Another studentwas concerned about evaluating the integrals in (23) and (24). I haven’t taught you how yet,but will eventually. The good news is that in contemporary research on the frontline fewtheoretical physicists actually have to calculate Feynman diagrams getting all the factorsof 2 right. In most cases, understanding the general behavior of the integral is sufficient.But of course, you should always take pride in getting everything right. In chapters II.6,III.6, and III.7 I will calculate Feynman diagrams for you in detail, getting all the factorsright so as to be able to compare with experiments.
Vacuum fluctuations
Let us go back to the terms I neglected in (18) and which you are supposed to havefigured out. For example, the four [ δ/δJ(w)]’s could have hit J
c,Jd,Jg, and Jhin (18)
thus producing something like
−iλ/integraldisplay/integraldisplay/integraldisplay/integraldisplay
DaeDbfJaJbJeJf(/integraldisplay
wDwwDww).
The coefficient of J(x1)J(x2)J(x3)J(x4)is then D13D24(−iλ/integraltext
wDwwDww)plus terms
obtained by permuting.
The corresponding physical process is easy to describe in words and in pictures (see
figure I.7.12). The source at x1produces a particle that propagates freely without any
interaction to x3, where it comes to a peaceful death. The particle produced at x2leads
a similarly uneventful life before being absorbed at x4. The two particles did not interact at
all. Somewhere off at the point w, which could perfectly well be anywhere in the universe,
there was an interaction with amplitude −iλ . This is known as a vacuum fluctuation:
t
xwx3 x4
x1 x2
Figure I.7.12
60 | I. Motivation and Foundation
As explained in chapter I.1, quantum mechanics and special relativity inevitably cause
particles to pop out of the vacuum, and they could even interact before vanishing againinto the vacuum. Look at different time slices (one of which is indicated by the dotted line)in figure I.7.12. In the far past, the universe has no particles. Then it has two particles,then four, then two again, and finally in the far future, none. We will have a lot more tosay in chapter VIII.2 about these fluctuations. Note that vacuum fluctuations occur also inour baby and child problems (see, e.g., Figs. I.7.1c, I.7.2g,h,i,j, and so forth).
Two words about history
I believe strongly that any self-respecting physicist should learn about the history of phys-ics, and the history of quantum field theory is among the most fascinating. Unfortunately,I do not have the space to go into it here.
4The path integral approach to field theory using
sources J(x) outlined here is mainly associated with Julian Schwinger, who referred to it
as “sorcery” during my graduate student days (so that I could tell people who inquired thatI was studying sorcery in graduate school.) In one of the many myths retold around tribalfires by physicists, Richard Feynman came upon his rules in a blinding flash of insight.In 1949 Freeman Dyson showed that the Feynman rules which so mystified people at thePocono conference a year earlier could actually be obtained from the more formal work ofJulian Schwinger and of Shin-Itiro Tomonaga.
Exercises
I.7.1 Work out the amplitude corresponding to figure I.7.11 in (24).
I.7.2 Derive (23) from first principles, that is from (11). It is a bit tedious, but straightforward. You should find
a symmetry factor1
2.
I.7.3 Draw all the diagrams describing two mesons producing four mesons up to and including order λ2.
Write down the corresponding Feynman amplitudes.
I.7.4 By Lorentz invariance we can always take k1+k2=(E,/vector0)in (23). The integral can be studied as a function
ofE. Show that for both internal particles to become real Emust be greater than 2 m. Interpret physically
what is happening.
4An excellent sketch of the history of quantum field theory is given in chapter 1 of The Quantum Theory of
Fields by S. Weinberg. For a fascinating history of Feynman diagrams, see Drawing Theories Apart by D. Kaiser.
I.8 Quantizing Canonically
Quantum electrodynamics is made to appear more difficult
than it actually is by the very many equivalent methods bywhich it may be formulated.
—R. P. Feynman
Always create before we annihilate, not the other way
around.
—Anonymous
Complementary formalisms
I adopted the path integral formalism as the quickest way to quantum field theory. But Imust also discuss the canonical formalism, not only because historically it was the methodused to develop quantum field theory, but also because of its continuing importance.Interestingly, the canonical and the path integral formalisms often appear complementary,in the sense that results difficult to see in one are clear in the other.
Heisenberg and Dirac
Let us begin with a lightning review of Heisenberg’s approach to quantum mechanics.
Given a classical Lagrangian for a single particle L=1
2˙q2−V( q) (we set the mass equal to
1), the canonical momentum is defined as p≡δL/δ˙q=˙q. The Hamiltonian is then given
byH=p˙q−L=p2/2+V( q) . Heisenberg promoted position q(t) and momentum p(t)
to operators by imposing the canonical commutation relation
[p,q]=−i (1)
Operators evolve in time according to
dp
dt=i[H,p]=−V/prime(q) (2)
62 | I. Motivation and Foundation
and
dq
dt=i[H,q]=p (3)
In other words, operators constructed out of pandqevolve according to O(t)=
eiHtO(0)e−iHt. In (1) pandqare understood to be taken at the same time. We obtain
the operator equation of motion ¨q=−V/prime(q) by combining (2) and (3).
Following Dirac, we invite ourselves to consider at some instant in time the operator
a≡(1/√
2ω)(ωq +ip) with some parameter ω. From (1) we have
[a,a†]=1 (4)
The operator a(t) evolves according to
da
dt=i[H,1√
2ω(ωq+ip)]=−i/radicalbiggω
2/parenleftbigg
ip+1
ωV/prime(q)/parenrightbigg
which can be written in terms of aanda†. The ground state |0/angbracketrightis defined as the state such
thata|0/angbracketright=0.
In the special case V/prime(q)=ω2qwe get the particularly simple resultda
dt=−iωa . This is
of course the harmonic oscillator L=1
2˙q2−1
2ω2q2andH=1
2(p2+ω2q2)=ω(a†a+1
2).
The generalization to many particles is immediate. Starting with
L=/summationdisplay
a1
2˙qa2−V( q1,q2,...,qN)
we have pa=δL/δ˙qaand
[pa(t),qb(t)]=−iδab (5)
The generalization to field theory is almost as immediate. In fact, we just use our handy-
dandy substitution table (I.3.6) and see that in D−dimensional space Lgeneralizes to
L=/integraldisplay
dDx{1
2(˙ϕ2−(/vector∇ϕ)2−m2ϕ2)−u(ϕ)} (6)
where we denote the anharmonic term (the interaction term in quantum field theory) by
u(ϕ) . The canonical momentum density conjugate to the field ϕ(/vectorx,t)is then
π(/vectorx,t)=δL
δ˙ϕ(/vectorx,t)=∂0ϕ(/vectorx,t) (7)
so that the canonical commutation relation at equal times now reads [using (I.3.6)]
[π(/vectorx,t),ϕ(/vectorx/prime,t)]=[∂0ϕ(/vectorx,t),ϕ(/vectorx/prime,t)]=−iδ(D)(/vectorx−/vectorx/prime) (8)
(and of course also [ π(/vectorx,t),π(/vectorx/prime,t)]=0 and [ ϕ(/vectorx,t),ϕ(/vectorx/prime,t)]=0.)Note that δabin (5)
gets promoted to δ(D)(/vectorx−/vectorx/prime)in (8). You should check that (8) has the correct dimension.
Turning the canonical crank we find the Hamiltonian
H=/integraldisplay
dDx[π(/vectorx,t)∂0ϕ(/vectorx,t)−L]
=/integraldisplay
dDx{1
2[π2+(/vector∇ϕ)2+m2ϕ2]+u(ϕ)} (9)
I.8. Quantizing Canonically | 63
For the case u=0, corresponding to the harmonic oscillator, we have a free scalar field
theory and can go considerably farther. The field equation reads
(∂2+m2)ϕ=0 (10)
Fourier expanding, we have
ϕ(/vectorx,t)=/integraldisplaydDk/radicalbig
(2π)D2ωk[a(/vectork)e−i(ω kt−/vectork./vectorx)+a†(/vectork)ei(ωkt−/vectork./vectorx)] (11)
withωk=+/radicalbig
/vectork2+m2so that the field equation (10) is satisfied. The peculiar looking factor
(2ωk)−1
2is chosen so that the canonical relation
[a(/vectork),a†(/vectork/prime)]=δ(D)(/vectork−/vectork/prime) (12)
appropriate for creation and annihilation operators implies the canonical commutation
[∂0ϕ(/vectorx,t),ϕ(/vectorx/prime,t)]=−iδ(D)(/vectorx−/vectorx/prime)in (8 ). You should check this but you can see why the
factor (2ωk)−1
2is needed since in ∂0ϕa factor ωkis brought down from the exponential.
As in quantum mechanics, the vacuum or ground state ||0/angbracketrightis defined by the condition
a(/vectork)|0/angbracketright=0 for all /vectorkand the single particle state by |/vectork/angbracketright≡a†(/vectork)|0/angbracketright. Thus, for example,
using (12) we have /angbracketleft0|ϕ(/vectorx,t)|/vectork/angbracketright=(1//radicalbig
(2π)D2ωk)e−i(ω kt−/vectork./vectorx), which we could think of
as the relativistic wave function of a single particle with momentum /vectork. For later use, we
will write this more compactly as (1/ρ(k))e−ik.x, with ρ(k)≡/radicalbig
(2π)D2ωka normalization
factor and k0=ωk.
To make contact with the path integral formalism let us calculate /angbracketleft0|ϕ(/vectorx,t)ϕ(/vector0, 0)|0/angbracketright
fort> 0. Of the four terms a†a†,a†a,aa†, and aa in the product of the two fields
onlyaa†survives, and thus using (12) we obtain/integraltext
[dDk/(2π)D2ωk]e−i(ω kt−/vectork./vectorx). In other
words, if we define the time-ordered product T[ϕ(x)ϕ(y) ]=θ(x0−y0)ϕ(x)ϕ(y) +θ(y0−
x0)ϕ(y)ϕ(x),w ef i n d
/angbracketleft0|T[ϕ(/vectorx,t)ϕ(/vector0, 0)]|0/angbracketright=
/integraldisplaydDk
(2π)D2ωk[θ(t)e−i(ω kt−/vectork./vectorx)+θ(−t)e+i(ω kt−/vectork./vectorx)] (13)
Go back to (I.3.23). We discover that /angbracketleft0|T[ϕ(x)ϕ( 0)]|0/angbracketright=iD(x) , the propagator for a
particle to go from 0 to xwe obtained using the path integral formalism. This further
justifies the iεprescription in (I.3.22). The physical meaning is that we always create
before we annihilate, not the other way around. This is a form of causality as formulatedin quantum field theory.
A remark: The combination d
Dk/(2ωk), even though it does not look Lorentz invariant,
is in fact a Lorentz invariant measure. To see this, we use (I.2.13) to show that (exercise I.8.1)
/integraldisplay
d(D+1)kδ(k2−m2)θ(k0)f (k0,/vectork)=/integraldisplaydDk
2ωkf( ωk,/vectork) (14)
for any function f( k) . Since Lorentz transformations cannot change the sign of k0, the step
function θ(k0)is Lorentz invariant. Thus the left-hand side is manifestly Lorentz invariant,
and hence the right-hand side must also be Lorentz invariant. This shows that relationssuch as (13) are Lorentz invariant; they are frame-independent statements.
64 | I. Motivation and Foundation
Scattering amplitude
Now that we have set up the canonical formalism it is instructive to see how the invariant
amplitude Mdefined in the preceding chapter arises using this alternative formalism. Let
us calculate the amplitude /angbracketleft/vectork3/vectork4|e−iHT|/vectork1/vectork2/angbracketright=/angbracketleft/vectork3/vectork4|ei/integraltext
d4xL(x)|/vectork1/vectork2/angbracketrightfor meson scatter-
ing/vectork1+/vectork2→/vectork3+/vectork4in order λwithu(ϕ)=λ
4!ϕ4. (We have, somewhat sloppily, turned
the large transition time Tinto/integraltext
dx0when going over to the Lagrangian.) Expanding in
λ, we obtain (−iλ
4!)/integraltext
d4x/angbracketleft/vectork3/vectork4|ϕ4(x)|/vectork1/vectork2/angbracketright.
The calculation of the matrix element is not dissimilar to the one we just did for the
propagator. There we have the product of two field operators between the vacuum state.Here we have the product of four field operators, all evaluated at the same spacetime pointx, sandwiched between two-particle states. There we look for a term of the form a(/vectork)a
†(/vectork),
Here, plugging the expansion (11) of the field into the product ϕ4(x), we look for terms of
the form a†(/vectork4)a†(/vectork3)a(/vectork2)a(/vectork1), so that we could remove the two incoming particles and
produce the two outgoing particles. (To avoid unnecessary complications we assume thatall four momenta are different.) We now annihilate and create. The annihilation operatora(/vectork
1)could have come from any one of the four ϕfields in ϕ4, giving a factor of 4, a(/vectork1)
could have come from any one of the three remaining ϕfields, giving a factor of 3, a†(/vectork3)
could have come from either of the two remaining ϕfields, giving a factor of 2, so that we
end up with a factor of 4!, which cancels the factor of1
4!included in the definition of λ.
(This is of course why, for the sake of convenience, λis defined the way it is. Recall from
the preceding chapter an analogous step.)
As you just learned and as you can see from (11), we obtain a factor of 1 /ρ(k)e−ik.xfor
each incoming particle and of 1 /ρ(k)e+ik.xfor each outgoing particle, giving all together
/parenleftbigg
/Pi14
α=11
ρ(kα)/parenrightbigg/integraldisplay
d4xei(k3+k 4−k 1−k 2).x=/parenleftbigg
/Pi14
α=11
ρ(kα)/parenrightbigg
(2π)4δ4(k3+k4−k1−k2)
It is conventional to refer to Sfi=/angbracketleftf|e−iHT|i/angbracketright, with some initial and final state, as
elements of an “S -matrix” and to define the “transition matrix” T-matrix by S=I+iT,
that is,
Sfi=δfi+iTfi (15)
In general, for initial and final states consisting of scalar particles characterized only by
their momenta, we write (using an obvious notation, for example/summationtext
ikis the sum of the
particle momenta in the initial state), invoking momentum conservation:
iTfi=(2π)4δ4⎛
⎝/summationdisplay
fk−/summationdisplay
ik⎞
⎠/parenleftbigg
/Pi1α1
ρ(kα)/parenrightbigg
M(f←i) (16)
In our simple example, iT(/vectork3/vectork4,/vectork1/vectork2)=(−iλ
4!)/integraltext
d4x/angbracketleft/vectork3/vectork4|ϕ4(x)|/vectork1/vectork2/angbracketright, and our little cal-
culation showed that M=−iλ, precisely as given in the preceding chapter. But this con-
nection between TfiandM, being “merely” kinematical, should hold in general, with the
invariant amplitude Mdetermined by the Feynman rules. I will not give a long boring for-
I.8. Quantizing Canonically | 65
mal proof, but you should check this assertion by working out some more involved cases,
such as the scattering amplitude to order λ2, recovering (I.7.23), for example.
Thus, quite pleasingly, we see that the invariant amplitude M determined by the
Feynman rules represents the “heart of the matter” with the momentum conservationdelta function and normalization factors stripped away.
My pedagogical aim here is merely to make one more contact (we will come across
more in later chapters) between the canonical and path integral formalisms, givingthe simplest possible example avoiding all subtleties and complications. Those readersinto rigor are invited to replace the plane wave states |/vectork
1/vectork2/angbracketrightwith wave packet states/integraltext
d3k1/integraltext
d3k2f1(/vectork1)f2(/vectork2)|/vectork1/vectork2/angbracketrightfor some appropriate functions f1andf2, starting in the
far past when the wave packets were far apart, evolving into the far future, so on and soforth, all the while smiling with self-satisfaction. The entire procedure is after all no differ-ent from the treatment of scattering
1in elementary nonrelativistic quantum mechanics.
Complex scalar field
Thus far, we have discussed a hermitean (often called real in a minor abuse of terminology)scalar field. Consider instead (as we will in chapter I.10) a nonhermitean (usually calledcomplex in another minor abuse) scalar field governed by L=∂ϕ
†∂ϕ−m2ϕ†ϕ.
Again, following Heisenberg, we find the canonical momentum density conjugate to
the field ϕ(/vectorx,t), namely π(/vectorx,t)=δL/ [δ˙ϕ(/vectorx,t)]=∂0ϕ†(/vectorx,t), so that [ π(/vectorx,t),ϕ(/vectorx/prime,t)]=
[∂0ϕ†(/vectorx,t),ϕ(/vectorx/prime,t)]=−iδ(D)(/vectorx−/vectorx/prime). Similarly, the canonical momentum density conju-
gate to the field ϕ†(/vectorx,t)is∂0ϕ(/vectorx,t).
Varying ϕ†we obtain the Euler-Lagrange equation of motion (∂2+m2)ϕ=0. [Similarly,
varying ϕwe obtain (∂2+m2)ϕ†=0.] Once again, we could Fourier expand, but now the
nonhermiticity of ϕmeans that we have to replace (11) by
ϕ(/vectorx,t)=/integraldisplaydDk/radicalbig
(2π)D2ωk/bracketleftBig
a(/vectork)e−i(ω kt−/vectork./vectorx)+b†(/vectork)ei(ωkt−/vectork./vectorx))/bracketrightBig
(17)
In (11) hermiticity fixed the second term in the square bracket in terms of the first. Here
in contrast, we are forced to introduce two independent sets of creation and annihilation
operators (a,a†)and(b,b†). You should verify that the canonical commutation relations
imply that these indeed behave like creation and annihilation operators.
Consider the current
Jμ=i(ϕ†∂μϕ−∂μϕ†ϕ) (18)
Using the equations of motion you should check that ∂μJμ=i(ϕ†∂2ϕ−∂2ϕ†ϕ). (This
follows immediately from the fact that the equation of motion for ϕ†is the hermitean
1For example, M. L. Goldberger and K. M. Watson, Collision Theory.
66 | I. Motivation and Foundation
conjugate of the equation of motion for ϕ.) The current is conserved and the corresponding
time-independent charge is given by (verify this!)
Q=/integraldisplay
dDxJ0(x)=/integraldisplay
dDk[a†(/vectork)a(/vectork)−b†(/vectork)b(/vectork)].
Thus the particle created by a†(call it the “particle”) and the particle created by b†(call it
the “antiparticle”) carry opposite charges. Explicitly, using the commutation relation wehaveQa
†|0/angbracketright=+a†|0/angbracketrightandQb†|0/angbracketright=−b†|0/angbracketright.
We conclude that ϕ†creates a particle and annihilates an antiparticle, that is, it produces
one unit of charge. The field ϕdoes the opposite. You should understand this point
thoroughly, as we will need it when we come to the Dirac field for the electron and positron.
The energy of the vacuum
As an instructive exercise let us calculate in the free scalar field theory the expectationvalue/angbracketleft0|H|0/angbracketright=/integraltext
d
Dx1
2/angbracketleft0|(π2+(/vector∇ϕ)2+m2ϕ2)|0/angbracketright, which we may loosely refer to as the
“energy of the vacuum.” It is merely a matter of putting together (7), (11), and (12). Let usfocus on the third term in /angbracketleft0|H|0/angbracketright, which in fact we already computed, since
/angbracketleft0|ϕ(/vectorx,t)ϕ(/vectorx,t)|0/angbracketright=/angbracketleft 0|ϕ(/vector0, 0)ϕ(/vector0, 0)|0/angbracketright
= lim
/vectorx,t→/vector0,0/angbracketleft0|ϕ(/vectorx,t)ϕ(/vector0, 0)|0/angbracketright= lim
/vectorx,t→/vector0,0/integraldisplaydDk
(2π)D2ωke−i(ω kt−/vectork./vectorx)=/integraldisplaydDk
(2π)D2ωk
The first equality follows from translation invariance, which also implies that the factor/integraltext
dDxin/angbracketleft0|H|0/angbracketrightcan be immediately replaced by V, the volume of space. The calculation
of the other two terms proceeds in much the same way: for example, the two factors of /vector∇
in(/vector∇ϕ)2just bring down a factor of /vectork2. Thus
/angbracketleft0|H|0/angbracketright=V/integraldisplaydDk
(2π)D2ωk1
2(ω2
k+/vectork2+m2)=V/integraldisplaydDk
(2π)D1
2/planckover2piωk (19)
upon restoring /planckover2pi.
You should find this result at once gratifying and alarming, gratifying because we recog-
nize it as the zero point energy of the harmonic oscillator integrated over all momentummodes and over all space, and alarming because the integral over /vectorkclearly diverges. But we
should not be alarmed: the energy of any physical configuration, for example the mass of aparticle, is to be measured relative to the “energy of the vacuum.” We ask for the differencein the energy of the world with and without the particle. In other words, we could simplydefine the correct Hamiltonian to be H−/angbracketleft0|H|0/angbracketright. We will come back to some of these
issues in chapters II.5, III.1, and VIII.2.
Nobody is perfect
In the canonical formalism, time is treated differently from space, and so one might worry
about the Lorentz invariance of the resulting field theory. In the standard treatment given
I.8. Quantizing Canonically | 67
in many texts, we would go on from this point and use the Hamiltonian to generate the
dynamics, developing a perturbation theory in the interaction u(ϕ). After a number of
formal steps, we would manage to derive the Feynman rules, which manifestly define aLorentz-invariant theory.
Historically, there was a time when people felt that quantum field theory should be
defined by its collection of Feynman rules, which gives us a concrete procedure to calculatemeasurable quantities, such as scattering cross sections. An extreme view along this lineheld that fields are merely mathematically crutches used to help us arrive at the Feynmanrules and should be thrown away at the end of the day.
This view became untenable starting in the 1970s, when it was realized that there
is a lot more to quantum field theory than Feynman diagrams. Field theory containsnonperturbative effects invisible by definition to Feynman diagrams. Many of these effects,which we will get to in due time, are more easily seen using the path integral formalism.
As I said, the canonical and the path integral formalism often appear to be complemen-
tary, and I will refrain from entering into a discussion about which formalism is superior.In this book, I adopt a pragmatic attitude and use whatever formalism happens to be easierfor the problem at hand.
Let me mention, however, some particularly troublesome features in each of the two
formalisms. In the canonical formalism fields are quantum operators containing an in-finite number of degrees of freedom, and sages once debated such delicate questions ashow products of fields are to be defined. On the other hand, in the path integral formal-ism, plenty of sins can be swept under the rug known as the integration measure (seechapter IV .7).
Appendix 1
It may seem a bit puzzling that in the canonical formalism the propagator has to be defined with time ordering,
which we did not need in the path integral formalism. To resolve this apparent puzzle, it suffices to look atquantum mechanics.
LetA[q(t
1)] be a function of q, evaluated at time t1. What does the path integral/integraltext
Dq(t) A [q(t 1)]ei/integraltextT
0dtL(˙q,q)
represent in the operator language? Well, working backward to (I.2.4) we see that we would slip A[q(t 1)] into the
factor /angbracketleftqj+1|e−iHδt|qj/angbracketright, where the integer jis determined by the condition that the time t1occurs between the
times jδt and(j+1)δt. In the resulting factor /angbracketleftqj+1|e−iHδtA[q(t1)]|qj/angbracketright, we could replace the c-number A[q(t1)]
by the operator A[ˆq], since A[ˆq]|qj/angbracketright=A[ qj]|qj/angbracketright/similarequalA[ q(t 1)]|qj/angbracketrightto the accuracy we are working with. Note that ˆqis
evidently a Schr ¨odinger operator. Thus, putting in this factor /angbracketleftqj+1|e−iHδtA[ˆq]|qj/angbracketrighttogether with all the factors of
/angbracketleftqi+1|e−iHδt|qi/angbracketright, we find that the integral in question, namely/integraltext
Dq(t) A [q(t 1)]ei/integraltextT
0dtL(˙q,q), actually represents
/angbracketleftqF|e−iH(T −t1)A[ˆq]e−iHt 1|qI/angbracketright=/angbracketleftqF|e−iHTA[ˆq(t1)]|qI/angbracketright
where ˆq(t 1)is now evidently a Heisenberg operator. [We have used the standard relation between Heisenberg
and Schr ¨odinger operators, namely, OH(t)=eiHtOSe−iHt.] I find this passage back and forth between the Dirac,
Schr ¨odinger, and Heisenberg pictures quite instructive, perhaps even amusing.
We are now prepared to ask the more complicated question: what does the path integral
/integraltext
Dq(t) A [q(t1)]B[q(t2)]ei/integraltextT
0dtL(˙q,q)represent in the operator language? Here B[q(t2)] is some other function of
qevaluated at time t2. So we also slip B[q(t 2)] into the appropriate factor in (I.2.4) and replace B[q(t2)]b yB[ˆq].
But now we see that we have to keep track of whether t1ort2is the earlier of the two times. If t2is earlier, the
operator A[ˆq] would appear to the left of the operator B[ˆq], and if t1is earlier, to the right. Explicitly, if t2is earlier
68 | I. Motivation and Foundation
thant1,we would end up with the sequence
e−iH(T −t1)A[ˆq]e−iH(t 1−t2)B[ˆq]e−iHt 2=e−iHTA[ˆq(t1)]B[ˆq(t2)] (20)
upon passing from the Schr ¨odinger to the Heisenberg picture, just as in the simpler situation above. Thus we
define the time-ordered product
T[A[ˆq(t1)]B[ˆq(t2)]]≡θ(t1−t2)A[ˆq(t1)]B[ˆq(t2)]+θ(t2−t1)B[ˆq(t2)]A[ˆq(t1)] (21)
We just learned that
/angbracketleftqF|e−iHTT[A[ˆq(t 1)]B[ˆq(t 2)]]|qI/angbracketright=/integraldisplay
Dq(t) A [q(t 1)]B[q(t 2)]ei/integraltextT
0dtL(˙q,q)(22)
The concept of time ordering does not appear on the right-hand side, but is essential on the left-hand side.
Generalizing the discussion here, we see that the Green’s functions G(n)(x1,x2,...,xn)introduced in the
preceding chapter [see (I.7.13–15)] is given in the canonical formalism by the vacuum expectation value of a time-ordered product of field operators /angbracketleft0|T{ϕ(x
1)ϕ(x2)...ϕ(xn)}|0/angbracketright. That (13) gives the propagator is a special case
of this relationship.
We could also consider /angbracketleft0|T{O 1(x1)O 2(x2)...On(xn)}|0/angbracketright, the vacuum expectation value of a time-ordered
product of various operators Oi(x) [the current Jμ(x), for example] made out of the quantum field. Such objects
will appear in later chapters [ for example, (VII.3.7)].
Appendix 2: Field redefinition
This is perhaps a good place to reveal to the innocent reader that there does not exist an international commissionin Brussels mandating what field one is required to use. If we use ϕ, some other guy is perfectly entitled to use
η, assuming that the two fields are related by some invertible function with η=f( ϕ) . (To be specific, it is often
helpful to think of η=ϕ+αϕ
3with some parameter α.) This is known as a field redefinition, an often useful
thing to do, as we will see repeatedly.
TheS-matrix amplitudes that experimentalists measure are invariant under field redefinition. But this is
tautological trivia: the scattering amplitude /angbracketleft/vectork3/vectork4|e−iHT|/vectork1/vectork2/angbracketright, for example, does not even know about ϕandη.
The issue is with the formalism we use to calculate the S-matrix.
In the path integral formalism, it is also trivial that we could write Z(J)=/integraltext
Dη ei[S(η)+/integraltext
d4xJη ]just as well
asZ(J)=/integraltext
Dϕ ei[S(ϕ)+/integraltext
d4xJϕ ]. This result, a mere change of integration variable, was known to Newton and
Leibniz. But suppose we write ˜Z(J)=/integraltext
Dϕ ei[S(ϕ)+/integraltext
d4xJη ]. Now of course any dolt could see that ˜Z(J)/negationslash=Z(J) ,
and a fortiori, the Green’s functions (I.7.14,15) obtained by differentiating ˜Z(J) andZ(J) are not equal.
The nontrivial physical statement is that the S-matrix amplitudes obtained from ˜Z(J) andZ(J) are in fact
the same. This better be the case, since we are claiming that the path integral formalism provides a way to actualphysics. To see how this apparent “miracle" (Green’s functions completely different, S-matrix amplitudes the
same) occurs, let us think physically. We set up our sources to produce or remove one single field disturbance,
as indicated in figure I.4.1. Our friend, who uses ˜Z(J) , in contrast, set up his sources to produce or remove
η=ϕ+αϕ
3(we specialize for pedagogical clarity), so that once in a while (with a probability determined by α)
he is producing three field disturbances instead of one, as shown in figure I.8.1. As a result, while he thinksthat he is scattering four mesons, occasionally he is actually scattering six mesons. (Perhaps he should give hisaccelerator a tune up.)
But to obtain S-matrix amplitudes we are told to multiply the Green’s functions by (k
2−m2)for each external
leg carrying momentum k, and then set k2tom2. When we do this, the diagram in figure I.8.1a survives, since
it has a pole that goes like 1 /(k2−m2)but the extraneous diagram in figure I.8.1b is eliminated. Very simple.
One point worth emphasizing is that mhere is the actual physical mass of the particle. Let’s be precise when we
should. Take the single particle state |/vectork/angbracketright. Act on it with the Hamiltonian. Then H|/vectork/angbracketright=/radicalbig
/vectork2+m2|/vectork/angbracketright. The mthat
appears in the eigenvalue of the Hamiltonian is the actual physical mass. We will come back to the issue of thephysical mass in chapter III.3.
In the canonical formalism, the field is an operator, and as we saw just now, the calculation of S-matrix
amplitudes involves evaluating products of field operators between physical states. In particular, the matrix
elements /angbracketleft/vectork|ϕ|0/angbracketrightand/angbracketleft0|ϕ|/vectork/angbracketright(related by hermitean conjugation) come in crucially. If we use some other field η,
I.8. Quantizing Canonically | 69
/H9272/H9272/H9272
/H9272
(a) (b)JJ
Figure I.8.1
what matters is merely that /angbracketleft/vectork|η|0/angbracketrightis not zero, in which case we could always write /angbracketleft/vectork|η|0/angbracketright=Z1
2/angbracketleft/vectork|ϕ|0/angbracketrightwith
Zsome c-number. We simply divide the scattering amplitude by the appropriate powers of Z1
2.
Exercises
I.8.1 Derive (14). Then verify explicitly that dDk/(2ωk)is indeed Lorentz invariant. Some authors prefer to
replace/radicalbig
2ωkin (11) by 2 ωkwhen relating the scalar field to the creation and annihilation operators.
Show that the operators defined by these authors are Lorentz covariant. Work out their commutationrelation.
I.8.2 Calculate /angbracketleft/vectork
/prime|H|/vectork/angbracketright, where |/vectork/angbracketright=a†(/vectork)|0/angbracketright.
I.8.3 For the complex scalar field discussed in the text calculate /angbracketleft0|T[ϕ(x)ϕ†(0)]|0/angbracketright.
I.8.4 Show that [ Q,ϕ(x) ]=−ϕ(x) .
I.9 Disturbing the Vacuum
Casimir effect
In the preceding chapter, we computed the energy of the vacuum /angbracketleft0|H|0/angbracketrightand obtained
the gratifying result that it is given by the zero point energy of the harmonic oscillatorintegrated over all momentum modes and over space. I explained that the energy of anyphysical configuration, for example, the mass of a particle, is to be measured relative tothis vacuum energy. In effect, we simply subtract off this vacuum energy and define thecorrect Hamiltonian to be H−/angbracketleft0|H|0/angbracketright.
But what if we disturb the vacuum?Physically, we could compare the energy of the vacuum before and after we introduce
the disturbance, by varying the disturbance for example. Of course, it is not just ourtextbook scalar field that contributes to the energy of the vacuum. The electromagneticfield, for instance, also undergoes quantum fluctuation and would contribute, with itstwo polarization degrees of freedom, to the energy density εof the vacuum the amount
2/integraltext
d
3k/(2π)31
2/planckover2piωk. In 1948 Casimir had the brilliant insight that we could disturb the
vacuum and produce a shift /Delta1ε. While εis not observable, /Delta1εshould be observable since
we can control how we disturb the vacuum. In particular, Casimir considered introducingtwo parallel “perfectly” conducting plates (formally of zero thickness and infinite extent)into the vacuum. The variation of /Delta1εwith the distance dbetween the plates would lead to
a force between the plates, known as the Casimir force. In reality, it is the electromagneticfield that is responsible, not our silly scalar field.
Call the direction perpendicular to the plates the xaxis. Because of the boundary
conditions the electromagnetic field must satisfy on the conducting plates, the wave vector
/vectorkcan only take on the values (πn/d ,k
y,kz), with nan integer. Thus the energy per unit
area between the plates is changed to/summationtext
n/integraltext
dkydkz/(2π)2/radicalBig
(πn/d)2+k2
y+k2
z.
To calculate the force, we vary d, but then we would have to worry about how the energy
density outside the two plates varies. A clever trick exists for avoiding this worry: weintroduce three plates! See figure (I.9.1). We hold the two outer plates fixed and move
I.9. Disturbing the Vacuum | 71
Ld L− d
Figure I.9.1
only the inner plate. Now we don’t have to worry about the world outside the plates. The
separation Lbetween the two outer plates can be taken to be as large as we like.
In the spirit of this book (and my philosophy) of avoiding computational tedium as
much as possible, I propose two simplifications: (I) do the calculation for a massless scalarfield instead of the electromagnetic field so we won’t have to worry about polarizationand stuff like that, and (II) retreat to a (1+1)-dimensional spacetime so we won’t have
to integrate over k
yandkz. Readers of quantum field theory texts do not need to watch
their authors show off their prowess in doing integrals. As you will see, the calculation isexceedingly instructive and gives us an early taste of the art of extracting finite physicalresults from seemingly infinite expressions, known as regularization, that we will studyin chapters III.1–3.
With this set-up, the energy E=f( d)+f( L−d)with
f( d)=1
2∞/summationdisplay
n=1ωn=π
2d∞/summationdisplay
n=1n (1)
since the modes are given by sin (nπx/d) (n =1,...,∞) with the corresponding energy
ωn=nπ/d .
Aagh! What do we do with∞/summationtext
n=1n? None of the ancient Greeks from Zeno on could tell us.
What they should tell us is that we are doing physics, not mathematics! Physical plates
cannot keep arbitrarily high frequency waves from leaking out.1
To incorporate this piece of all-important physics, we should introduce a factor e−aωn/π
with a parameter ahaving the dimension of time (or in our natural units, length) so
that modes with ωn/greatermuchπ/a do not contribute: they don’t see the plates! The characteristic
1See footnote 1 in chapter III.1.
72 | I. Motivation and Foundation
frequency π/a parametrizes the high frequency response of the conducting plates. Thus
we have
f( d)=π
2d∞/summationdisplay
n=1ne−an/d=−π
2∂
∂a∞/summationdisplay
n=1e−an/d=−π
2∂
∂a1
1−e−a/d=π
2dea/d
(ea/d−1)2
Since we want a−1to be large, we take the limit asmall so that
f( d)=πd
2a2−π
24d+πa2
480d3+O(a4/d5). (2)
Note that f( d) blows up as a→0, as it should, since we are then back to (1). But the force
between two conducting plates shouldn’t blow up. Experimentalists might have noticed itby now!
Well, the force is given by
F=−∂E
∂d=− {f/prime(d)−f/prime(L−d)}=−/braceleftbigg/parenleftbigg1
2πa2+π
24d2+.../parenrightbigg
−(d→L−d)/bracerightbigg
−→
a→0−π
24/parenleftbigg1
d2−1
(L−d)2/parenrightbigg
−→
L/greatermuchd−π/planckover2pic
24d2(3)
Behold, the parameter awe have introduced to make sense of the divergent sum in (1) has
disappeared in our final result for the physically measurable force. In the last step, usingdimensional analysis we restored /planckover2pito underline the quantum nature of the force.
The Casimir force between two plates is attractive. Notice that the 1 /d
2of the force
simply follows from dimensional analysis since in natural units force has dimension ofan inverse length squared. In a tour de force, experimentalists have measured this tinyforce. The fluctuating quantum field is quite real!
To obtain a sensible result we need to regularize in the ultraviolet (namely the high
frequency or short time behavior parametrized by a) and in the infrared (namely the long
distance cutoff represented by L). Notice how aandL“work” together in (3).
This calculation foreshadows the renormalization of quantum field theories, a topic
made mysterious and almost incomprehensible in many older texts. In fact, it is perfectlysensible. We will discuss renormalization in detail in chapters III.1 and III.2, but for nowlet us review what we just did.
Instead of panicking when faced with the divergent sum in (1), we remind ourselves
that we are proud physicists and that physics tells us that the sum should not go all theway to infinity. In a conducting plate, electrons rush about to counteract any applied tan-gential electric field. But when the incident wave oscillates at sufficiently high frequency,the electrons can’t keep up. Thus the idealization of a perfectly conducting plate fails. Weregularize (such an ugly term but that’s what field theorists use!) the sum in a mathemati-cally convenient way by introducing a damping factor. The single parameter ais supposed
to summarize the unknown high frequency physics that causes the electron to fail to keepup. In reality, a
−1is related to the plasma frequency of the metal making up the plate.
I.9. Disturbing the Vacuum | 73
A priori, the Casimir force between the two plates could end up depending on the
parameter a. In that case, the Casimir force would tell us something about the response of a
conducting plate to high frequency electric fields, and it would have made for an interestingchapter in a text on solid state physics. Since this is in fact a quantum field theory text, youmight have suspected that the Casimir force represents some fundamental physics aboutfluctuating quantum fields and that awould drop out. That the Casimir force is “universal”
makes it unusually interesting. Notice however, as is physically sensible, that the O(1/d
4)
correction to the Casimir force does depend on whether the experimentalist used copperor aluminum plates.
We might then wonder whether the leading term F=−π/(24d
2)depends on the
particular regularization we used. What if we suppress the higher terms in the divergentsum with some other function? We will address this question in the appendix to thischapter.
Amusingly, the 24 in (3) is the same 24 that appears in string theory! (The dimension
of spacetime the quantum bosonic string must live in is determined to be 24 +2=26.)
The reader who knows string theory would know what these two cryptic statements areabout (summing up the zero modes of the string). Appallingly, in an apparent attempt tomake the subject appear even more mysterious than it is, some treatments simply assert
that the sum
∞/summationtext
n=1nis by some mathematical sleight-of-hand equal to −1/12. Even though
it would have allowed us to wormhole from (1) to (3) instantly, this assertion is manifestly
absurd. What we did here, however, makes physical sense.
Appendix
Here we address the fundamental issue of whether a physical quantity we extract by cutting off high frequency
contributions could depend on how we cut. Let me mention that in recent years the study of Casimir force foractual physical situations has grown into an active area of research, but clearly my aim here is not to give a realisticaccount of this field, but to study in an easily understood context an issue (as you will see) central to quantumfield theory. My hope is that by the time you get to actually regularize a field theory in (3+1)-dimensional
spacetime, you would have amply mastered the essential physics and not have to struggle with the mechanics ofregularization.
Let us first generalize a bit the regularization scheme we used and write
f( d)=π
2d∞/summationdisplay
n=1ng/parenleftbiggna
d/parenrightbigg
=π
2∂
∂a∞/summationdisplay
n=1h/parenleftbiggna
d/parenrightbigg
≡π
2∂
∂aH/parenleftbigga
d/parenrightbigg
(4)
Hereg(v)=h/prime(v) is a rapidly decreasing function so that the sums make sense, chosen judiciously to allow
ready evaluation of H(a/d) ≡∞/summationtext
n=1h(na/d). [In (2), we chose g(v)=e−vand hence h(v)=−e−v.] We would like
to know how the Casimir force,
−F=∂f (d)
∂d−(d→L−d)=π
2∂2
∂d∂aH(a
d)−(d→L−d) (5)
depends on g(v) .
Let us try to get as far as we can using physical arguments and dimensional analysis. Expand Has follows:
πH(a/d) =...+γ−2d2/a2+γ−1d/a+γ0+γ1a/d+γ2a2/d2+.... We might be tempted to just write a Taylor
series in a/d , but nothing tells us that H(a/d) might not blow up as a→0. Indeed, the example in (3) contains
a term like d/a , and so we better be cautious.
74 | I. Motivation and Foundation
We will presently argue physically that the series in fact terminate in one direction. The force is given by
F=/parenleftbigg
...+γ−22d
a3+γ−11
2a2+γ11
2d2+γ22a
d3+.../parenrightbigg
−(d→L−d) (6)
Look at the γ−2term: it contributes to the force a term like (d−(L−d))/a3. But as remarked earlier, the two
outer plates could be taken as far apart as we like. The force could not depend on L, and thus on physical grounds
γ−2must vanish. Similarly, all γ−kfork> 2 must vanish.
Next, we note that the γ−1, although definitely not zero, gets subtracted away since it does not depend on d.
(Theγ0term has already gone away.) At this point, notice that, furthermore, the γkterms with k> 2 all vanish
asa→0. You could check that all these assertions hold for the g(v) used in the text.
Remarkably, the Casimir force is determined by γ1alone: F=γ1/(2d2). As noted earlier, the force has to be
proportional to 1 /d2. This fact alone shows us that in (6) we only need to keep the γ1term. In the text, we found
γ1=−π/12. In exercise I.9.1 I invite you to go through an amusing calculation obtaining the same value for γ1
with an entirely different choice of g(v) .
This already suggests that the Casimir force is regularization independent, that it tells us more about the
vacuum than about metallic conductivity, but still it is highly instructive to study an entire class of damping or
regularizing functions to watch how regularization independence emerges. Let us regularize the sum over zero
point energies to f( d)=1
2∞/summationtext
n=1ωnK(ωn)with
K(ω) =/summationdisplay
αcα/Lambda1α
ω+/Lambda1α(7)
Herecαis a bunch of real numbers and /Lambda1α(known as regulators or regulator frequencies) a bunch of high
frequencies subject to certain conditions but otherwise chosen entirely at our discretion. For the sum∞/summationtext
n=1ωnK(ωn)
to converge, we need K(ωn)to vanish faster than 1 /ω2
n. In fact, for ωmuch larger than /Lambda1α,K(ω)→1
ω/summationtext
αcα/Lambda1α−
1
ω2/summationtext
αcα/Lambda1α2+.... The requirement that the 1 /ωand 1/ω2terms vanish gives the conditions
/summationdisplay
αcα/Lambda1α=0 (8)
and
/summationdisplay
αcα/Lambda12
α=0 (9)
respectively.
Furthermore, low frequency physics is not to be modified, and so we want K(ω) →1 forω< </Lambda1α, thus
requiring
/summationdisplay
αcα=1 (10)
At this point, we do not even have to specify the set the index αruns over beyond the fact that the three conditions
(8),(9), and (10) require that αmust take on at least three values. Note also that some of the cα’s must be negative.
Incidentally, we could do with fewer regulators if we are willing to invoke some knowledge of metals, for instance,thatK(ω) =K(−ω) , but that is not the issue here.
We now show that the Casimir force between the two plates does not depend on the choices of c
αand/Lambda1α.
First, being physicists rather than mathematicians, we freely interchange the two sums in f( d) and write
f( d)=1
2/summationdisplay
αcα/Lambda1α/summationdisplay
nωn
ωn+/Lambda1α=−1
2/summationdisplay
αcα/Lambda1α/summationdisplay
n/Lambda1α
ωn+/Lambda1α(11)
I.9. Disturbing the Vacuum | 75
where, without further ceremony, we have used condition (8). Next, keeping in mind that the sum/summationtext
n/Lambda1α/(ωn+/Lambda1α)is to be put back into (11), we massage it (defining for convenience bα=π//Lambda1 α) as follows:
∞/summationdisplay
n=1/Lambda1
ωn+/Lambda1=∞/summationdisplay
n=1/integraldisplay∞
0dte−t( 1+nb
d)=/integraldisplay∞
0dte−t/parenleftBigg
1
1−e−bt
d−1/parenrightBigg
=/integraldisplay∞
0dte−t[d
tb−1
2+tb
12d+O/parenleftBig
b3/parenrightBig
] (12)
(To avoid clutter we have temporarily suppressed the index α.) All these manipulations make perfect sense since
the entire expression is to be inserted into the sum over αin (11) after we restore the index α. It appears that the
result would depend on cαandλα. In fact, mentally restoring and inserting, we see that the 1 /bterm in (12) can
be thrown away since
/summationdisplay
αcα/Lambda1α/bα=π/summationdisplay
αcα/Lambda12
α=0 (13)
[There is in fact a bit of an overkill here since this term corresponds to the γ−1term, which does not appear in
the force anyway. Thus the condition (9) is, strictly speaking, not necessary. We are regularizing not merely the
force, but f( d) so that it defines a sensible function.] Similarly, the b0term in (12) can be thrown away thanks
to (8). Thus, keeping only the bterm in (12), we obtain
f( d)=−1
24d/integraldisplay∞
0dte−tt/summationdisplay
αcα/Lambda1αbα+O/parenleftbigg1
d3/parenrightbigg
=−π
24d+O/parenleftbigg1
d3/parenrightbigg
(14)
Indeed, f( d) , and a fortiori the Casimir force, do not depend on the cα’s and /Lambda1α’s. To the level of rigor enter-
tained by physicists (but certainly not mathematicians), this amounts to a proof of regularization independencesince with enough regulators we could approximate any (reasonable) function K(ω) that actually describes real
conducting plates. Again, as is physically sensible, you could check that the O(1/d
3)term in f( d) does depend
on the regularization scheme.
The reason that I did this calculation in detail is that we will encounter this class of regularization, known as
Pauli-Villars, in chapter III.1 and especially in the calculation of the anomalous magnetic moment of the electronin chapter III.7, and it is instructive to see how regularization works in a more physical context before dealingwith all the complications of relativistic field theory.
Exercises
I.9.1 Choose the damping function g(v)=1/(1+v)2instead of the one in the text. Show that this re-
sults in the same Casimir force. [Hint: To sum the resulting series, pass to an integral representation
H(ξ)=−∞/summationtext
n=11/(1+nξ)=−∞/summationtext
n=1/integraltext∞
0dte−(1+nξ)t=/integraltext∞
0dte−t/(1−eξt). Note that the integral blows up
logarithmically near the lower limit, as expected.]
I.9.2 Show that with the regularization used in the appendix, the 1 /dexpansion of the force between two
conducting plates contains only even powers.
I.9.3 Show off your skill in doing integrals by calculating the Casimir force in (3+1)-dimensional spacetime.
For help, see M. Kardar and R. Golestanian, Rev. Mod. Phys. 71: 1233, 1999; J. Feinberg, A. Mann, and
M. Revzen, Ann. Phys. 288: 103, 2001.
I.10 Symmetry
Symmetry, transformation, and invariance
The importance of symmetry in modern physics cannot be overstated.1
When a law of physics does not change upon some transformation, that law is said to
exhibit a symmetry.
I have already used Lorentz invariance to restrict the form of an action. Lorentz invari-
ance is of course a symmetry of spacetime, but fields can also transform in what is thoughtof as an internal space. Indeed, we have already seen a simple example of this. I noted inpassing in chapter I.3 that we could require the action of a scalar field theory to be invariantunder the transformation ϕ→−ϕand so exclude terms of odd power in ϕ, such as ϕ
3,
from the action.
With the ϕ3term included, two mesons could scatter and go into three mesons, for
example by the diagrams in (fig. I.10.1). But with this term excluded, you can easilyconvince yourself that this process is no longer allowed. You will not be able to drawa Feynman diagram with an odd number of external lines. (Think about modifying the
integral in our baby problem in chapter I.7 to/integraltext
+∞
−∞dqe−1
2m2q2−gq3−λq4+Jq.)Thus the
simple reflection symmetry ϕ→−ϕimplies that in any scattering process the number
of mesons is conserved modulo 2.
Now that we understand one scalar field, let us consider a theory with two scalar fields
ϕ1andϕ2satisfying the reflection symmetry ϕa→−ϕa(a=1o r2):
L=1
2(∂ϕ1)2−1
2m2
1ϕ2
1−λ1
4ϕ4
1+1
2(∂ϕ2)2−1
2m2
2ϕ2
2−λ2
4ϕ4
2−ρ
2ϕ2
1ϕ2
2(1)
We have two scalar particles, call them 1 and 2, with mass m1andm2. To lowest order,
they scatter in the processes 1 +1→1+1, 2+2→2+2, 1+2→1+2, 1+1→2+2,
and 2 +2→1+1 (convince yourself). With the five parameters m1,m2,λ1,λ2, and ρ
completely arbitrary, there is no relationship between the two particles.
1A. Zee, Fearful Symmetry .
I.10. Symmetry | 77
Figure I.10.1
It is almost an article of faith among theoretical physicists, enunciated forcefully by
Einstein among others, that the fundamental laws should be orderly and simple, ratherthan arbitrary and complicated. This orderliness is reflected in the symmetry of the action.
Suppose that m
1=m2andλ1=λ2; then the two particles would have the same mass and
their interaction, with themselves and with each other, would be the same. The LagrangianLbecomes invariant under the interchange symmetry ϕ
1←→ϕ2.
Next, suppose we further impose the condition ρ=λ1=λ2so that the Lagrangian
becomes
L=1
2/bracketleftBig
(∂ϕ 1)2+(∂ϕ 2)2/bracketrightBig
−1
2m2/parenleftBig
ϕ2
1+ϕ2
2/parenrightBig
−λ
4/parenleftBig
ϕ2
1+ϕ2
2/parenrightBig2
(2)
It is now invariant under the 2-dimensional rotation {ϕ1(x)→cosθϕ1(x)+sinθϕ2(x),
ϕ2(x)→−sin θϕ 1(x)+cosθϕ 2(x)} withθan arbitrary angle. We say that the theory
enjoys an “internal” SO( 2)symmetry, internal in the sense that the transformation has
nothing to do with spacetime. In contrast to the interchange symmetry ϕ1←→ϕ2the
transformation depends on the continuous parameter θ, and the corresponding symmetry
is said to be continuous.
We see from this simple example that symmetries exist in hierarchies.
Continuous symmetries
If we stare at the equations of motion (∂2+m2)ϕa=−λ/vectorϕ2ϕalong enough we see that if
we define Jμ≡i(ϕ1∂μϕ2−ϕ2∂μϕ1), then ∂μJμ=i(ϕ1∂2ϕ2−ϕ2∂2ϕ1)=0 so that Jμis a
conserved current. The corresponding charge Q=/integraltext
dDxJ0, just like electric charge, is
conserved.
Historically, when Heisenberg noticed that the mass of the newly discovered neutron
was almost the same as the mass of a proton, he proposed that if electromagnetism weresomehow turned off there would be an internal symmetry transforming a proton into aneutron.
An internal symmetry restricts the form of the theory, just as Lorentz invariance restricts
the form of the theory. Generalizing our simple example, we could construct a field theorycontaining Nscalar fields ϕ
a, with a=1,...,Nsuch that the theory is invariant under the
transformations ϕa→Rabϕb(repeated indices summed), where the matrix Ris an element
of the rotation group SO(N) (see appendix B for a review of group theory). The fields ϕa
transform as a vector /vectorϕ=(ϕ1,...,ϕN). We can form only one basic invariant, namely the
78 | I. Motivation and Foundation
cd
a b/H110022i/H9261 (/H9254ab/H9254cd /H11001 /H9254ac/H9254bd /H11001 /H9254ad/H9254bc)
Figure I.10.2
scalar product /vectorϕ./vectorϕ=ϕaϕa=/vectorϕ2(as always, repeated indices are summed unless otherwise
specified). The Lagrangian is thus restricted to have the form
L=1
2/bracketleftBig
(∂/vectorϕ)2−m2/vectorϕ2/bracketrightBig
−λ
4(/vectorϕ2)2(3)
The Feynman rules are given in fig. I.10.2. When we draw Feynman diagrams, each line
carries an internal index in addition to momentum.
Symmetry manifests itself in physical amplitudes. For example, imagine calculating
the propagator iDab(x)=/integraltext
D/vectorϕeiSϕa(x)ϕb(0). We assume that the measure D/vectorϕis in-
variant under SO(N). By thinking about how Dab(x) transforms under the symmetry
group SO(N) you see easily (exercise I.10.2) that it must be proportional to δab. You can
check this by drawing a few Feynman diagrams or by considering an ordinary integral/integraltext
d/vectorqe−S(q)qaqb. No matter how complicated a diagram you draw (fig. I.10.3, e.g.) you al-
ways get this factor of δab. Similarly, scattering amplitudes must also exhibit the symmetry.
Without the SO(N) symmetry, many other terms would be possible (e.g., ϕaϕbϕcϕdfor
some arbitrary choice of a,b,c, andd)in (3).
We can write R=eθ.Twhere θ.T=/summationtext
AθATAis a real antisymmetric matrix. The group
SO(N) hasN(N−1)/2 generators, which we denote by TA. [Think of the familiar case of
SO( 3).] Under an infinitesimal transformation (repeated indices summed) ϕa→Rabϕb/similarequal
(1+θATA)abϕb, or in other words, we have the infinitesimal change δϕa=θATAabϕb.
Noether’s theorem
We now come to one of the most profound observations in theoretical physics, namely
Noether’s theorem, which states that a conserved current is associated with each generator
ab
dhg
c
ef
ijm
kl
Figure I.10.3
I.10. Symmetry | 79
of a continuous symmetry. The appearance of a conserved current for (2) is not an acci-
dent.
As is often the case with the most important theorems, the proof of Noether’s theorem is
astonishingly simple. Denote the fields in our theory generically by ϕa.Since the symmetry
is continuous, we can consider an infinitesimal change δϕa. Since Ldoes not change,
we have
0=δL=δL
δϕaδϕa+δL
δ∂μϕaδ∂μϕa
=δL
δϕaδϕa+δL
δ∂μϕa∂μδϕa (4)
We would have been stuck at this point, but if we use the equations of motion δL/δϕa=
∂μ(δL/δ∂μϕa)we can combine the two terms and obtain
0=∂μ/parenleftBigg
δL
δ∂μϕaδϕa/parenrightBigg
(5)
If we define
Jμ≡δL
δ∂μϕaδϕa (6)
then (5) says that ∂μJμ=0. We have found a conserved current! [It is clear from the
derivation that the repeated index ais summed in (6)].
Let us immediately illustrate with the simple scalar field theory in (3). Plugging δϕa=
θA(TA)abϕbinto (6) and noting that θAis arbitrary, we obtain N(N−1)/2 conserved
currents JA
μ=∂μϕa(TA)abϕb, one for each generator of the symmetry group SO(N).
In the special case N=2, we can define a complex field ϕ≡(ϕ1+iϕ2)/√
2. The La-
grangian in (3) can be written as
L=∂ϕ†∂ϕ−m2ϕ†ϕ−λ(ϕ†ϕ)2,
and is clearly invariant under ϕ→eiθϕandϕ†→e−iθϕ†. We find from (6) that Jμ=
i(ϕ†∂μϕ−∂μϕ†ϕ), the current we met already in chapter I.8. Mathematically, this is
because the groups SO( 2)andU(1)are isomorphic (see appendix B).
For pedagogical clarity I have used the example of scalar fields transforming as a vector
under the group SO(N) . Obviously, the preceding discussion holds for an arbitrary group
Gwith the fields ϕtransforming under an arbitrary representation RofG. The conserved
currents are still given by JA
μ=∂μϕa(TA)abϕbwithTAthe generators evaluated in the
representation R. For example, if ϕtransform as the 5-dimensional representation of
SO( 3)thenTAi sa5b y5matrix.
For physics to be invariant under a group of transformations it is only necessary that the
action be invariant. The Lagrangian density Lcould very well change by a total divergence:
δL=∂μKμ, provided that the relevant boundary term could be dropped. Then we would
see immediately from (5) that all we have to do to obtain a formula for the conservedcurrent is to modify (6) to J
μ≡(δL/δ∂μϕa)δϕa−Kμ. As we will see in chapter VIII.4,
many supersymmetric field theories are of this type.
80 | I. Motivation and Foundation
Charge as generators
Using the canonical formalism of chapter I.8, we can derive an elegant result for the charge
associated with the conserved current
Q≡/integraldisplay
d3xJ0=/integraldisplay
d3xδL
δ∂0ϕaδϕa
Note that Qdoes not depend on the time at which the integral is evaluated:
dQ
dt=/integraldisplay
d3x∂0J0=−/integraldisplay
d3x∂iJi=0 (7)
Recognizing that δL/δ∂ 0ϕais just the canonical momentum conjugate to the field ϕa,w e
see that
i[Q,ϕa]=δϕa (8)
The charge operator generates the corresponding transformation on the fields. An impor-
tant special case is for the complex field ϕinSO( 2)/similarequalU(1)theory we discussed; then
[Q,ϕ]=ϕandeiθQϕe−iθQ=eiθϕ.
Exercises
I.10.1 Some authors prefer the following more elaborate formulation of Noether’s theorem. Suppose that the
action does not change under an infinitesimal transformation δϕa(x)=θAVA
a[withθAsome parameters
labeled by AandVA
asome function of the fields ϕb(x) and possibly also of their first derivatives with
respect to x]. It is important to emphasize that when we say the action Sdoes not change we are not
allowed to use the equations of motion. After all, the Euler-Lagrange equations of motion follow fromdemanding that δS=0 for any variation δϕ
asubject to certain boundary conditions. Our scalar field
theory example nicely illustrates this point, which is confused in some books: δS=0 merely because S
is constructed using the scalar product of O(N) vectors.
Now let us do something apparently a bit strange. Let us consider the infinitesimal change written
above but with the parameters θAdependent on x. In other words, we now consider δϕa(x)=θA(x)VA
a.
Then of course there is no reason for δSto vanish; but, on the other hand, we know that since δSdoes
vanish when θAis constant, δSmust have the form δS=/integraltext
d4xJμ(x)∂μθA(x). In practice, this gives us
a quick way of reading off the current Jμ(x); it is just the coefficient of ∂μθA(x) inδS.
Show how all this works for the Lagrangian in (3).
I.10.2 Show that Dab(x) must be proportional to δabas stated in the text.
I.10.3 Write the Lagrangian for an SO( 3)invariant theory containing a Lorentz scalar field ϕtransforming
in the 5-dimensional representation up to quartic terms. [Hint: It is convenient to write ϕas a 3 by 3
symmetric traceless matrix.]
I.10.4 Add a Lorentz scalar field ηtransforming as a vector under SO( 3)to the Lagrangian in exercise I.10.3,
maintaining SO( 3)invariance. Determine the Noether currents in this theory. Using the equations of
motion, check that the currents are conserved.
I.11 Field Theory in Curved Spacetime
General coordinate transformation
In Einstein’s theory of gravity, the invariant Minkowskian spacetime interval ds2=
ημνdxμdxν=(dt)2−(d/vectorx)2is replaced by ds2=gμνdxμdxν, where the metric tensor
gμν(x) is a function of the spacetime coordinates x. The guiding principle, known as the
principle of general covariance, states that physics, as embodied in the action S, must be
invariant under arbitrary coordinate transformations x→x/prime(x). More precisely, the prin-
ciple1states that with suitable restrictions the effect of a gravitational field is equivalent to
that of a coordinate transformation.
Since
ds2=g/prime
λσdx/primeλdx/primeσ=g/prime
λσ∂x/primeλ
∂xμ∂x/primeσ
∂xνdxμdxν=gμνdxμdxν
the metric transforms as
g/prime
λσ(x/prime)∂x/primeλ
∂xμ∂x/primeσ
∂xν=gμν(x) (1)
The inverse of the metric gμνis defined by gμνgνρ=δμ
ρ.
A scalar field by its very name does not transform: ϕ(x)=ϕ/prime(x/prime). The gradient of the
scalar field transforms as
∂μϕ(x)=∂x/primeλ
∂xμ∂ϕ/prime(x/prime)
∂x/primeλ=∂x/primeλ
∂xμ∂/prime
λϕ/prime(x/prime)
By definition, a (covariant) vector field transforms as
Aμ(x)=∂x/primeλ
∂xμA/prime
λ(x/prime)
1For a precise statement of the principle of general covariance, see S. Weinberg, Gravitation and Cosmology ,
p. 92.
82 | I. Motivation and Foundation
so that ∂μϕ(x) is a vector field. Given two vector fields Aμ(x) andBν(x), we can contract
them with gμν(x) to form gμν(x)Aμ(x)Bν(x), which, as you can immediately check, is
a scalar. In particular, gμν(x)∂μϕ(x)∂ νϕ(x) is a scalar. Thus, if we simply replace the
Minkowski metric ημνin the Lagrangian L=1
2[(∂ϕ)2−m2ϕ2]=1
2(ημν∂μϕ∂νϕ−m2ϕ2)by
the Einstein metric gμν, the Lagrangian is invariant under coordinate transformation.
The action is obtained by integrating the Lagrangian over spacetime. Under a coordinate
transformation d4x/prime=d4xdet(∂x/prime/∂x) . Taking the determinant of (1), we have
g≡detgμν=detg/prime
λσ∂x/primeλ
∂xμ∂x/primeσ
∂xν=g/prime/bracketleftbigg
det/parenleftbigg∂x/prime
∂x/parenrightbigg/bracketrightbigg2
(2)
We see that the combination d4x√−g=d4x/prime/radicalbig
−g/primeis invariant under coordinate transfor-
mation.
Thus, given a quantum field theory we can immediately write down the theory in curved
spacetime. All we have to do is promote the Minkowski metric ημνin our Lagrangian to the
Einstein metric gμνand include a factor√−g in the spacetime integration measure.2The
action Swould then be invariant under arbitrary coordinate transformations. For example,
the action for a scalar field in curved spacetime is simply
S=/integraldisplay
d4x√−g1
2(gμν∂μϕ∂νϕ−m2ϕ2) (3)
(There is a slight subtlety involving the spin1
2field that we will talk about in part II. We
will eventually come to it in chapter VIII.1.)
There is no essential difficulty in quantizing the scalar field in curved spacetime. We
simply treat gμνas given [e.g., the Schwarzschild metric in spherical coordinates: g00=
(1−2GM/r) ,grr=−(1−2GM/r)−1,gθθ=−r2, and gφφ=−r2sin2θ] and study the
path integral/integraltext
DϕeiS, which is still a Gaussian integral and thus do-able. The propagator
of the scalar field D(x ,y)can be worked out and so on and so forth. (see exercise I.11.1).
At this point, aside from the fact that gμν(x)carries Lorentz indices while ϕ(x) does not,
the metric gμνlooks just like a field and is in fact a classical field. Write the action of the
world S=Sg+SMas the sum of two terms: Sgdescribing the dynamics of the gravitational
fieldgμν(which we will come to in chapter VIII.1) and SMdescribing the dynamics of all
the other fields in the world [the “matter fields,” namely ϕin our simple example with SM
as given in (3)]. We could quantize gravity by integrating over gμνas well, thus extending
the path integral to/integraltext
DgDϕeiS.
Easier said than done! As you have surely heard, all attempts to integrate over gμν(x)
have been beset with difficulties, eventually driving theorists to seek solace in string theory.I will explain in due time why Einstein’s theory is known as “nonrenormalizable.”
2We also have to replace ordinary derivatives ∂μby the covariant derivatives Dμof general relativity, but acting
on a scalar field ϕthe covariant derivative is just the ordinary derivative Dμϕ=∂μϕ.
I.11. Field Theory in Spacetime | 83
What the graviton listens to
One of the most profound results to come out of Einstein’s theory of gravity is a funda-
mental definition of energy and momentum. What exactly are energy and momentum anyway? Energy and momentum are what the graviton listens to. (The graviton is of coursethe particle associated with the field g
μν.)
The stress-energy tensor Tμνis defined as the variation of the matter action SMwith
respect to the metric gμν(holding the coordinates xμfixed):
Tμν(x)=−2√−gδSM
δgμν(x)(4)
Energy is defined as E=P0=/integraltext
d3x√−gT00(x)and momentum as Pi=/integraltext
d3x√−gT0i(x).
Even if we are not interested in curved spacetime per se, (4) still offers us a simple (and
fundamental) way to determine the stress energy of a field theory in flat spacetime. Wesimply vary around the Minkowski metric η
μνby writing gμν=ημν+hμνand expand SM
to first order in h. According to (4), we have3
SM(h)=SM(h=0)−/integraldisplay
d4x/bracketleftBig
1
2hμνTμν+O(h2)/bracketrightBig
. (5)
The symmetric tensor field hμν(x) is in fact the graviton field (see chapters I.5 and
VIII.1). The stress-energy tensor Tμν(x) is what the graviton field couples to, just as the
electromagnetic current Jμ(x) is what the photon field couples to.
Consider a general SM=/integraltext
d4x√−g(A +gμνBμν+gμνgλρCμνλρ+...). Since −g=
1+ημνhμν+O(h2)andgμν=ημν−hμν+O(h2), we find
Tμν=2(Bμν+2Cμνλρηλρ+...)−ημνL (6)
in flat spacetime. Note
T≡ημνTμν=−(4A+2ημνBμν+0.ημνηλρCμνλρ+...) (7)
which we have written in a form emphasizing that Cμνλρ does not contribute to the trace
of the stress-energy tensor.
We now show the power of this definition of Tμνby obtaining long-familiar results
about the electromagnetic field. Promoting the Lagrangian of the massive spin 1 fieldto curved spacetime we have
4L=(−1
4gμνgλρFμλFνρ+1
2m2gμνAμAν)and thus5Tμν=
−FμλFλ
ν+m2AμAν−ημνL.
3I use the normal convention in which indices are summed regardless of any symmetry; in other words,
1
2hμνTμν=1
2(h01T01+h10T10+...)=h01T01+....
4Here we use the fact that the covariant curl is equal to the ordinary curl DμAν−DνAμ=∂μAν−∂νAμand
soFμνdoes not involve the metric.
5Holding xμfixed means that we hold ∂μand hence Aμfixed since Aμis related to ∂μby gauge invariance.
We are anticipating (see chapter II.7) here, but you have surely heard of gauge invariance in a nonrelativistic
context.
84 | I. Motivation and Foundation
For the electromagnetic field we set m=0. First, L=−1
4FμνFμν=−1
4(−2F2
0i+F2
ij)=
1
2(/vectorE2−/vectorB2). Thus, T00=−F0λFλ
0−1
2(/vectorE2−/vectorB2)=1
2(/vectorE2+/vectorB2). That was comforting, to
see a result we knew from “childhood.” Incidentally, this also makes clear that we canthink of /vectorE
2as kinetic energy and /vectorB2as potential energy. Next, T0i=−F0λFλ
i=F0jFij=
εijkEjBk=(/vectorE×/vectorB)i. The Poynting vector has just emerged.
Since the Maxwell Lagrangian L=−1
4gμνgλρFμλFνρinvolves only the Cterm with
Cμνλρ=−1
4FμνFλρ, we see from (7) that the stress-energy tensor of the electromagnetic
field is traceless, an important fact we will need in chapter VIII.1. We can of course checkdirectly that T=0 (exercise I.11.4).
6
Appendix: A concise introduction to curved spacetime
General relativity is often made to seem more difficult and mysterious than need be. Here I give a concise review
of some of its basic elements for later use.
Denote the spacetime coordinates of a point particle by Xμ. To construct its action note that the only
coordinate invariant quantity is the “length”7of the world line traced out by the particle (fig. I.11.1), namely/integraltext
ds=/integraltext/radicalbiggμνdXμdXν, where gμνis evaluated at the position of the particle of course. Thus, the action for a
point particle must be proportional to
/integraldisplay
ds=/integraldisplay/radicalBig
gμνdXμdXν=/integraldisplay/radicalBigg
gμν[X(ζ) ]dXμ
dζdXν
dζdζ
where ζis any parameter that varies monotonically along the world line. The length, being geometric, is
manifestly reparametrization invariant, that is, independent of our choice of ζas long as it is reasonable. This
is one of those “more obvious than obvious” facts since/integraltext/radicalbiggμνdXμdXνis manifestly independent of ζ.I fw e
insist, we can check the reparametrization invariance of/integraltext
ds. Obviously the powers of dζmatch. Explicitly, if
we write ζ=ζ(η) , then dXμ/dζ=(dη/dζ)(dXμ/dη) anddζ=(dζ/dη)dη .
Let us define
K≡gμν[X(ζ) ]dXμ
dζdXν
dζ
for ease of writing. Setting the variation of/integraltext
dζ√
Kequal to zero, we obtain
/integraldisplay
dζ1√
K(2gμνdXμ
dζdδXν
dζ+∂λgμνdXμ
dζdXν
dζδXλ)=0
which upon integration by parts (and with δXλ=0 at the endpoints as usual) gives the equation of motion
√
Kd
dζ/parenleftbigg1√
K2gμλdXμ
dζ/parenrightbigg
−∂λgμνdXμ
dζdXν
dζ=0 (8)
To simplify (8) we exploit our freedom in choosing ζand set dζ=ds, so that K=1. We have
2gμλd2Xμ
ds2+2∂σgμλdXσ
dsdXμ
ds−∂λgμνdXμ
dsdXν
ds=0
6We see that tracelessness is related to the fact that the electromagnetic field has no mass scale. Pure
electromagnetism is said to be scale or dilatation invariant. For more on dilatation invariance see S. Coleman,
Aspects of Symmetry ,p .6 7 .
7We put “length” in quotes because if gμνhad a Euclidean signature then/integraltext
dswould indeed be the length
and minimizing/integraltext
dswould give the shortest path (the geodesic) between the endpoints, but here gμνhas a
Minkowskian signature.
I.11. Field Theory in Spacetime | 85
Xμ(ζ)
Figure I.11.1
which upon multiplication by gρλbecomes
d2Xρ
ds2+1
2gρλ(2∂νgμλ−∂λgμν)dXμ
dsdXν
ds=0
that is,
d2Xρ
ds2+/Gamma1ρ
μν[X(s) ]dXμ
dsdXν
ds=0 (9)
if we define the Riemann-Christoffel symbol by
/Gamma1ρ
μν≡1
2gρλ(∂μgνλ+∂νgμλ−∂λgμν) (10)
Given the initial position Xμ(s0)and velocity (dXμ/ds)(s 0)we have four second order differential equations (9)
determining the geodesic followed by the particle in curved spacetime. Note that, contrary to the impressiongiven by some texts, unlike (8), (9) is not reparametrization invariant.
To recover Newtonian gravity, three conditions must be met: (1) the particle moves slowly dX
i/ds/lessmuchdX0/ds;
(2) the gravitational field is weak, so that the metric is almost Minkowskian gμν/similarequalημν+hμν; and (3) the
gravitational field does not depend on time. Condition (1) means that d2Xρ/ds2+/Gamma1ρ
00(dX0/ds)2/similarequal0, while (2)
and (3) imply that /Gamma1ρ
00/similarequal−1
2ηρλ∂λh00. Thus, (9) reduces to d2X0/ds2/similarequal0 (which implies that dX0/ds is a constant)
andd2Xi/ds2+1
2∂ih00(dX0/ds)2/similarequal0, which since X0is proportional to sbecomes d2Xi/dt2/similarequal−1
2∂ih00. Thus,
we obtain Newton’s equationd2/vectorX
dt2/similarequal−/vector∇φif we define the gravitational potential φbyh00=2φ:
g00/similarequal1+2φ (11)
Referring to the Schwarzschild metric, we see that far from a massive body, φ=−GM/r , as we expect. (Note
also that this derivation depends neither on hijnor on h0j, as long as they are time independent.)
Thus, the action of a point particle is
S=−m/integraldisplay/radicalBig
gμνdXμdXν=−m/integraldisplay/radicalBigg
gμν[X(ζ) ]dXμ
dζdXν
dζdζ (12)
Themfollows from dimensional analysis.
A slick way of deriving S(which also allows us to see the minus sign) is to start with the nonrelativistic action
of a particle in a gravitational potential φ, namely S=/integraltext
Ldt=/integraltext
(1
2mv2−m−mφ)dt . Note that the rest mass
86 | I. Motivation and Foundation
mcomes in with a minus sign as it is part of the potential energy in nonrelativistic physics. Now force Sinto a
relativistic form: For small vandφ,
S=−m/integraldisplay
(1−1
2v2+φ)dt/similarequal−m/integraldisplay/radicalbig
1−v2+2φdt
=−m/integraldisplay/radicalbig
(1+2φ)(dt)2−(d/vectorx)2
We see8that the 2 in (11) comes from the square root in the Lorentz-Fitzgerald quantity√
1−v2.
Now that we have Swe can calculate the stress energy of a point particle using (4):
Tμν(x)=m√−g/integraldisplay
dζK−1
2δ(4)[x−X(ζ) ]dXμ
dζdXν
dζ
Setting ζtos(which we call the proper time τin this context) we have
Tμν(x)=m√−g/integraldisplay
dτδ(4)[x−X(τ) ]dXμ
dτdXν
dτ
In particular, as we expect the 4-momentum of the particle is given by
Pν=/integraldisplay
d3x√−gT0ν=m/integraldisplay
dτδ [x0−X0(τ)]dX0
dτdXν
dτ=mdXν
dτ
The action (12) given here has two defects: (1) it is difficult to deal with a path integral
/integraltext
DXe−im/integraltext
dζ√
(dXμ/dζ)(dX μ/dζ)involving a square root, and (2) Sdoes not make sense for a massless particle.
To remedy these defects, note that classically, Sis equivalent to
Simp=−1
2/integraldisplay
dζ/parenleftbigg1
γdXμ
dζdXμ
dζ+γm2/parenrightbigg
(13)
where (dXμ/dζ)(dXμ/dζ)=gμν(X)(dXμ/dζ)(dXν/dζ). Varying with respect to γ(ζ) we obtain m2γ2=
(dXμ/dζ)(dXμ/dζ) . Eliminating γinSimp we recover S.
The path integral/integraltext
DXeiSimphas a standard quadratic form.9Quantum mechanics of a relativistic point parti-
cle is best formulated in terms of Simp, notS. Furthermore, for m=0,Simp=−1
2/integraltext
dζ[γ−1(dXμ/dζ)(dX μ/dζ)]
makes perfect sense. Note that varying with respect to γnow gives the well-known fact that for a massless particle
gμν(X)dXμdXν=0.
The action Simp will provide the starting point for our discussion on string theory in chapter VIII.5.
Exercises
I.11.1 Integrate by parts to obtain for the scalar field action
S=−/integraldisplay
d4x√−g1
2ϕ(1√−g∂μ√−ggμν∂ν+m2)ϕ
8We should not conclude from this that gij=δij.The point is that to leading order in v/c, our particle is
sensitive only to g00, as we have just shown. Indeed, restoring cin the Schwarzschild metric we have
ds2=(1−2GM
c2r)c2dt2−(1−2GM
c2r)−1dr2−r2dθ2−r2sin2θdφ2
→c2dt2−d/vectorx2−2GM
rdt2+O(1/c2)
9One technical problem, which we will address in chapter III.4, is that in the integral over X(ζ) apparently
different functions X(ζ) may in fact be the same physically, related by a reparametrization.
I.11. Field Theory in Spacetime | 87
and write the equation of motion for ϕin curved spacetime. Discuss the propagator of the scalar field
D(x ,y)(which is of course no longer translation invariant, i.e., it is no longer a function of x−y).
I.11.2 Use (4) to find Efor a scalar field theory in flat spacetime. Show that the result agrees with what you
would obtain using the canonical formalism of chapter I.8.
I.11.3 Show that in flat spacetime Pμas derived here from the stress energy tensor Tμνwhen interpreted as
an operator in the canonical formalism satisfies [ Pμ,ϕ(x) ]=−i∂μϕ(x) , and thus does exactly what you
expect the energy and momentum operators to do, namely to be conjugate to time and space and hencerepresented by −i∂
μ.
I.11.4 Show that for the Maxwell field Tij=−(EiEj+BiBj)+1
2δij(/vectorE2+/vectorB2)and hence T=0.
I.12 Field Theory Redux
What have you learned so far?
Now that we have reached the end of part I, let us take stock of what you have learned.
Quantum field theory is not that difficult; it just consists of doing one great big integral
Z(J)=/integraldisplay
Dϕ ei/integraltext
dD+1x[1
2(∂ϕ)2−1
2m2ϕ2−λϕ4+Jϕ ](1)
By repeatedly functionally differentiating Z(J) and then setting J=0 we obtain
/integraldisplay
Dϕ ϕ(x 1)ϕ(x 2)...ϕ(xn)ei/integraltext
dD+1x[1
2(∂ϕ)2−1
2m2ϕ2−λϕ4](2)
which tells us about the amplitude for nparticles associated with the field ϕto come into
and go out of existence at the spacetime points x1,x2,...,xn, interacting with each other
in between. Birth and death, with some kind of life in between.
Ah, if we could only do the integral in (1)! But we can’t. So one way of going about it is
to evaluate the integral as a series in λ:
∞/summationdisplay
k=0(−iλ)k
k!/integraldisplay
Dϕ ϕ(x 1)ϕ(x 2)...ϕ(xn)[/integraldisplay
dD+1yϕ(y)4]kei/integraltext
dD+1x[1
2(∂ϕ)2−1
2m2ϕ2]
(3)
To keep track of the terms in the series we draw little diagrams.
Quantum field theorists try to dream up ways to evaluate (1), and failing that, they invent
tricks and methods for extracting the physics they are interested in, by hook and by crook,without actually evaluating (1).
To see that quantum field theory is a straightforward generalization of quantum me-
chanics, look at how (1) reduces appropriately. We have written the theory in (D+1)-
dimensional spacetime, that is, Dspatial dimensions and 1 temporal dimension. Consider
(1) in(0+1)-dimensional spacetime, that is, no space; it becomes
Z(J)=/integraldisplay
Dϕ ei/integraltext
dt[1
2(dϕ
dt)2−1
2m2ϕ2−λϕ4+Jϕ ](4)
I.12. Field Theory Redux | 89
where we now denote the spacetime coordinate xjust by time t. We recognize this as the
quantum mechanics of an anharmonic oscillator with the position of the mass point tiedto the spring denoted by ϕand with an external force Jpushing on the oscillator.
In the quantum field theory (1), each term in the action makes physical sense: The first
two terms generalize the harmonic oscillator to include spatial variations, the third termthe anharmonicity, and the last term an external probe. You can think of a quantum fieldtheory as an infinite collection of anharmonic oscillators, one at each point in space.
We have here a scalar field ϕ. In previous and future chapters, the notion of field was
and will be generalized ever so slightly: The field can transform according to a nontrivialrepresentation of the Lorentz group. We have already encountered fields transforming asa vector and a tensor and will presently encounter a field transforming as a spinor. Lorentzinvariance and whatever other symmetries we have constrain the form of the action. Theintegral will look more complicated but the approach is exactly as outlined here.
That’s just about all there is to quantum field theory.
This page intentionally left blank
Part II Dirac and the Spinor
This page intentionally left blank
II.1 The Dirac Equation
Staring into a fire
According to a physics legend, apparently even true, Dirac was staring into a fire one
evening in 1928 when he realized that he wanted, for reasons that are no longer relevant,a relativistic wave equation linear in spacetime derivatives ∂
μ≡∂/∂xμ. At that time, the
Klein-Gordon equation (∂2+m2)ϕ=0, which describes a free particle of mass mand
quadratic in spacetime derivatives, was already well known. This is in fact the equation ofmotion of the scalar field theory we studied earlier.
At first sight, what Dirac wanted does not make sense. The equation is supposed to
have the form “some linear combination of ∂
μacting on some field ψis equal to some
constant times the field.” Denote the linear combination by cμ∂μ. If the cμ’s are four
ordinary numbers, then the four-vector cμdefines some direction and the equation cannot
be Lorentz invariant.
Nevertheless, let us follow Dirac and write, using modern notation,
(iγμ∂μ−m)ψ=0 (1)
At this point, the four quantities iγμare just the coefficients of ∂μandmis just a constant.
We have already argued that γμcannot simply be four numbers. Well, let us see what these
objects have to be in order for this equation to contain the correct physics.
Acting on (1) with (iγμ∂μ+m), Dirac obtained −(γμγν∂μ∂ν+m2)ψ=0. It is tradi-
tional to define, in addition to the commutator [ A,B]=AB−BA familiar from quan-
tum mechanics, the anticommutator {A,B}=AB +BA. Since derivatives commute,
γμγν∂μ∂ν=1
2{γμ,γν}∂μ∂ν, and we have (1
2{γμ,γν}∂μ∂ν+m2)ψ=0. In a moment of
inspiration Dirac realized that if
{γμ,γν}=2ημν(2)
withημνthe Minkowski metric he would obtain (∂2+m2)ψ=0, which describes a particle
of mass m, and thus (1) would also describe a particle of mass m.
94 | II. Dirac and the Spinor
Since ημνis a diagonal matrix with diagonal elements η00=1 and ηjj=− 1, (2) says
that(γ0)2=1,(γj)2=− 1, and γμγν=−γνγμforμ/negationslash=ν. This last statement, that the
coefficients γμanticommute with each other, implies that they indeed cannot be ordinary
numbers. Dirac’s thought would make sense if we could find four such objects.
Clifford algebra
A set of objects γμ(clearly dof them in d-dimensional spacetime) satisfying the relation
(2) is said to form a Clifford algebra. I will develop the mathematics of Clifford algebralater. Suffice it for you to check here that the following 4 by 4 matrices satisfy (2):
γ0=/parenleftBiggI 0
0−I/parenrightBigg
=I⊗τ3 (3)
γi=/parenleftBigg0 σi
−σi0/parenrightBigg
=σi⊗iτ2 (4)
Hereσandτdenote the standard Pauli matrices. For historical reasons the four matrices
γμare known as gamma matrices—not a very imaginative name! (Our convention is such
that whether an index on a Pauli matrix is upper or lower has no significance. On the otherhand, we define γ
μ≡ημνγνand it does matter whether the index on a gamma matrix is
upper or lower; it is to be treated just like the index on any Lorentz vector. This conventionis useful because then γ
μ∂μ=γμ∂μ.)
The direct product notation is convenient for computation: For example, γiγj=(σi⊗
iτ2)(σj⊗iτ2)=(σiσj⊗i2τ2τ2)=−(σiσj⊗I)and thus {γi,γj}=− { σi,σj}⊗I=
−2δijas desired.
You can convince yourself that the γμ’s cannot be smaller than 4 by 4 matrices. The
mathematics forces the Dirac spinor ψto have 4 components! The physical content of
the Dirac equation (1) is most transparent if we transform to momentum space: we plugψ(x)=/integraltext
[d
4p/(2π)4]e−ipxψ(p) into (1) and obtain
(γμpμ−m)ψ(p) =0 (5)
Since (5) is Lorentz invariant, as we will show below, we can examine its physical content
in any frame, in particular the rest frame pμ=(m,/vector0), in which it becomes
(γ0−1)ψ=0 (6)
As(γ0−1)2=− 2(γ0−1)we recognize (γ0−1)as a projection operator up to a trivial nor-
malization. Indeed, using the explicit form in (3), we see that there is nothing mysterious
to Dirac’s equation: When written out, (6) reads
/parenleftBigg00
0I/parenrightBigg
ψ=0
thus telling us that 2 of the 4 components in ψare zero.
II.1. The Dirac Equation | 95
This makes perfect sense since we know that the electron has 2 physical degrees of
freedom, not 4. Viewed in this light, the mysterious Dirac equation is no more and no lessthan a projection that gets rid of the unwanted degrees of freedom. Compare our discussionof the equation of motion of a massive spin 1 particle (chapter I.5). There also, 1 of the 4components of A
μis projected out. Indeed, the Klein-Gordon equation (∂2+m2)ϕ(x) =0
just projects out those Fourier components ϕ(k) not satisfying the mass shell condition
k2=m2. Our discussion provides a unified view of the equations of motion in relativistic
physics: They just project out the unphysical components.
A convenient notation introduced by Feynman, /negationslasha≡γμaμfor any 4-vector aμ, is now
standard. The Dirac equation then reads (i/negationslash∂−m)ψ=0.
Cousins of the gamma matrices
Under a Lorentz transformation x/primeν=/Lambda1ν
μxμ, the 4 components of the vector field Aμ
transform like, well, a vector. How do the 4 components of ψtransform? Surely not in the
same way as Aμsince even under rotation ψandAμtransform quite differently: one as
spin1
2and the other as spin 1. Let us write ψ(x)→ψ/prime(x/prime)≡S(/Lambda1)ψ(x) and try to determine
the 4 by 4 matrix S(/Lambda1) .
It is a good idea to first sort out (and name) the 16 linearly independent 4 by 4 matrices.
We already know five of them: the identity matrix and the γμ’s. The strategy is simply
to multiply the γμ’s together, thus generating more 4 by 4 matrices until we get all 16.
Since the square of a gamma matrix γμis equal to ±1 and the γμ’s anticommute with
each other, we have to consider only γμγν,γμγνγλ, andγμγνγλγρwithμ,ν,λ, andρall
different from one another. Thus, the only product of four gamma matrices that we haveto consider is
γ5≡iγ0γ1γ2γ3(7)
This combination is so important that it has its own name! (The peculiar name comes
about because in some old-fashioned notation the time coordinate was called x4with a
corresponding γ4.)We have
γ5=i(I⊗τ3)(σ1⊗iτ2)(σ2⊗iτ2)(σ3⊗iτ2)=i4(I⊗τ3)(σ1σ2σ3⊗τ2)
and so
γ5=I⊗τ1=/parenleftBigg0I
I 0/parenrightBigg
(8)
With the factor of iincluded, γ5is manifestly hermitean. An important property is that γ5
anticommutes with the γμ’s:
{γ5,γμ}=0 (9)
96 | II. Dirac and the Spinor
Continuing, we see that the products of three gamma matrices, all different, can be
written as γμγ5(e.g.,γ1γ2γ3=−iγ0γ5). Finally, using (2) we can write the product of
two gamma matrices as γμγν=ημν−iσμν, where
σμν≡i
2[γμ,γν] (10)
There are 4 .3/2=6 of these σμνmatrices.
Count them, we got all 16. The set of 16 matrices {1, γμ,σμν,γμγ5,γ5}forms a
complete basis of the space of all 4 by 4 matrices, that is, any 4 by 4 matrix can be writtenas a linear combination of these 16 matrices.
It is instructive to write out σ
μνexplicitly in the representation (3) and (4):
σ0i=i/parenleftBigg0σi
σi0/parenrightBigg
(11)
σij=εijk/parenleftBiggσk0
0σk/parenrightBigg
(12)
We see that σijare just the Pauli matrices doubly stacked, for example,
σ12=/parenleftBiggσ30
0σ3/parenrightBigg
Lorentz transformation
Recall from a course on quantum mechanics that a general rotation can be written as
ei/vectorθ/vectorJwith/vectorJthe 3 generators of rotation and /vectorθ3 rotation parameters. Recall also that the
Lorentz group contains boosts in addition to rotations, with /vectorKdenoting the 3 generators of
boosts. Recall from a course on electromagnetism that the 6 generators {/vectorJ,/vectorK}transform
under the Lorentz group as the components of an antisymmetric tensor just like theelectromagnetic field F
μνand thus can be denoted by Jμν. I will discuss these matters
in more detail in chapter II.3. For the moment, suffice it to note that with this notation
we can write a Lorentz transformation as /Lambda1=e−i
2ωμνJμν, with Jijgenerating rotations,
J0igenerating boosts, and the antisymmetric tensor ωμν=−ωνμwith its 6 =4.3/2
components corresponding to the 3 rotation and 3 boost parameters.
Given the preceding discussion and the fact that there are six matrices σμν, we suspect
that up to an overall numerical factor the σμν’s must represent the 6 generators Jμνof the
Lorentz group acting on a spinor. In fact, our suspicion is confirmed by thinking about
what a rotation e−i
2ωijJijdoes. Referring to (12) we see that if Jijis represented by1
2σij
this would correspond exactly to how a spin1
2particle transforms in quantum mechanics.
More precisely, separate the 4 components of the Dirac spinor into 2 sets of 2 components:
ψ=/parenleftBiggφ
χ/parenrightBigg
(13)
II.1. The Dirac Equation | 97
From (12) we see that under a rotation around the 3rd axis, φ→e−iω 121
2σ3φandχ→
e−iω 121
2σ3χ. It is gratifying to see that φandχtransform like 2-component Pauli spinors.
We have thus figured out that a Lorentz transformation /Lambda1acting on ψis represented
byS(/Lambda1)=e−(i/4 )ωμνσμν, and so, acting on ψ, the generators Jμνare indeed represented
by1
2σμν. Therefore we would expect that if ψ(x) satisfies the Dirac equation (1) then
ψ/prime(x/prime)≡S(/Lambda1)ψ(x) would satisfy the Dirac equation in the primed frame,
(iγμ∂/prime
μ−m)ψ/prime(x/prime)=0 (14)
where ∂/prime
μ≡∂/∂x/primeμ. To show this, calculate [ σμν,γλ]=2i(γμηνλ−γνημλ)and hence
forωinfinitesimal SγλS−1=γλ−(i/4)ωμν[σμν,γλ]=γλ+γμωλ
μ. Building up a finite
Lorentz transformation by compounding infinitesimal transformations (just as in thestandard discussion of the rotation group in quantum mechanics), we have Sγ
λS−1=
/Lambda1λ
μγμ.
Dirac bilinears
The Clifford algebra tells us that (γ0)2=+ 1 and (γi)2=− 1; hence the necessity for the i
in (4). One consequence of the iis that γ0is hermitean while γiis antihermitean, a fact
conveniently expressed as
(γμ)†=γ0γμγ0(15)
Thus, contrary to what you might think, the bilinear ψ†γμψis not hermitean; rather,
¯ψγμψis hermitean with ¯ψ≡ψ†γ0. The necessity for introducing ¯ψin addition to ψ†in
relativistic physics is traced back to the (+,−,−,−)signature of the Minkowski metric.
It follows that (σμν)†=γ0σμνγ0. Hence, S(/Lambda1)†=γ0e(i/4)ωμνσμνγ0, (which incidentally,
clearly shows that Sis not unitary, a fact we knew since σ0iis not hermitean), and so
¯ψ/prime(x/prime)=ψ(x)†S(/Lambda1)†γ0=¯ψ(x)e+(i/4 )ωμνσμν. (16)
We have
¯ψ/prime(x/prime)ψ/prime(x/prime)=¯ψ(x)e+(i/4 )ωμνσμνe−(i/4 )ωμνσμνψ(x)=¯ψ(x)ψ(x)
You are probably used to writing ψ†ψin nonrelativistic physics. In relativistic physics you
have to get used to writing ¯ψψ .I ti s ¯ψψ , notψ†ψ, that transforms as a Lorentz scalar.
There are obviously 16 Dirac bilinears ¯ψ/Gamma1ψ that we can form, corresponding to the 16
linearly independent /Gamma1’s. You can now work out how various fermion bilinears transform
(exercise II.1.1). The notation is rather nice: Various objects transform the way it looks like
they should transform. We simply look at the Lorentz indices they carry. Thus, ¯ψ(x)γμψ(x)
transforms as a Lorentz vector.
98 | II. Dirac and the Spinor
Parity
An important discrete symmetry in physics is that of parity or reflection in a mirror1
xμ→x/primeμ=(x0,−/vectorx) (17)
Multiply the Dirac equation (1) by γ0:γ0(iγμ∂μ−m)ψ(x) =0=(iγμ∂/prime
μ−m)γ0ψ(x) ,
where ∂/prime
μ≡∂/∂x/primeμ. Thus,
ψ/prime(x/prime)≡ηγ0ψ(x) (18)
satisfies the Dirac equation in the space-reflected world (where ηis an arbitrary phase that
we can set to 1).
Note, for example, ¯ψ/prime(x/prime)ψ/prime(x/prime)=¯ψ(x)ψ(x) but¯ψ/prime(x/prime)γ5ψ/prime(x/prime)=¯ψ(x)γ0γ5γ0ψ(x)=
−¯ψ(x)γ5ψ(x) . Under a Lorentz transformation ¯ψ(x)γ5ψ(x) and¯ψ(x)ψ(x) transform in
the same way but under space reflection they transform in an opposite way; in other words,while ¯ψ(x)ψ(x) transforms as a scalar, ¯ψ(x)γ
5ψ(x) transforms as a pseudoscalar.
You are now ready to do the all-important exercises in this chapter.
The Dirac Lagrangian
An interesting question: What Lagrangian would give Dirac’s equation? The answer is
L=¯ψ(i/negationslash∂−m)ψ (19)
Since ψis complex we can vary ψand¯ψindependently to obtain the Euler-Lagrange
equation of motion. Thus, ∂μ(δL/δ∂μψ)−δL/δψ=0 gives ∂μ(i¯ψγμ)+m¯ψ=0, which
upon hermitean conjugation and multiplication by γ0gives the Dirac equation (1). The
other variational equation ∂μ(δL/δ∂μ¯ψ)−δL/δ¯ψ=0 gives the Dirac equation even more
directly. (If you are disturbed by the asymmetric treatment of ψand¯ψ, you can always
integrate by parts in the action, have ∂μact on ¯ψin the Lagrangian, and then average the
two forms of the Lagrangian. The action S=/integraltext
d4xLtreats ψand¯ψsymmetrically.)
Slow and fast electrons
Given a set of gamma matrices it is straightforward to solve the Dirac equation
(/negationslashp−m)ψ(p) =0 (20)
forψ(p) : It is a simple matrix equation (see exercise II.1.3).
1Rotations consist of all linear transformations xi→Rijxjsuch that det R=+ 1. Those transformations with
detR=− 1 are composed of parity followed by a rotation. In (3+1)-dimensional spacetime, parity can be defined
as reversing one of the spatial coordinates or all three spatial coordinates. The two operations are related by arotation. Note that in odd dimensional spacetime, parity is not the same as space inversion, in which all spatial
coordinates are reversed (see exercise II.1.12).
II.1. The Dirac Equation | 99
Note that if somebody uses the gamma matrices γμ, you are free to use instead γ/primeμ=
W−1γμWwithWany 4 by 4 matrix with an inverse. Obviously, γ/primeμalso satisfy the Clifford
algebra. This freedom of choice corresponds to a simple change of basis. Physics cannotdepend on the choice of basis, but which basis is the most convenient depends on thephysics.
For example, suppose we want to study a slowly moving electron. Let us use the basis
defined by (3) and (4), and the 2-component decomposition of ψ(13). Since (6) tells us
thatχ(p)=0 for an electron at rest, we expect χ(p) to be much smaller than φ(p) for a
slowly moving electron.
In contrast, for momentum much larger than the mass, we can approximate (20) by
/negationslashpψ(p) =0. Multiplying on the left by γ
5, we see that if ψ(p) is a solution then γ5ψ(p)
is also a solution since γ5anticommutes with γμ. Since (γ5)2=1, we can form two
projection operators PL≡1
2(1−γ5)andPR≡1
2(1+γ5), satisfying P2
L=PL,P2
R=PR,
andPLPR=0. It is extremely useful to introduce the two combinations ψL=1
2(1−γ5)ψ
andψR=1
2(1+γ5)ψ. Note that γ5ψL=−ψLandγ5ψR=+ψR. Physically, a relativistic
electron has two degrees of freedom known as helicities: it can spin either clockwise oranticlockwise around the direction of motion. I leave it to you as an exercise to show thatψ
LandψRcorrespond to precisely these two possibilities. The subscripts LandRindicate
left and right handed. Thus, for fast moving electrons, a basis known as the Weyl basis,designed so that γ
5, rather than γ0, is diagonal, is more convenient. Instead of (3), we
choose
γ0=/parenleftBigg0I
I 0/parenrightBigg
=I⊗τ1 (21)
We keep γias in (4). This defines the Weyl basis. We now calculate
γ5≡iγ0γ1γ2γ3=i(I⊗τ1)(σ1σ2σ3⊗i3τ2)=−(I⊗τ3)=/parenleftBigg−I 0
0I/parenrightBigg
(22)
which is indeed diagonal as desired. The decomposition into left and right handed fields is
of course defined regardless of what basis we feel like using, but in the Weyl basis we havethe nice feature that ψ
Lhas two upper components and ψRhas two lower components.
The spinors ψLandψRare known as Weyl spinors.
Note that in going from the Dirac to the Weyl basis γ0andγ5trade places (up to a sign):
Dirac : γ0diagonal; Weyl : γ5diagonal. (23)
Physics dictates which basis to use: We prefer to have γ0diagonal when we deal with
slowly moving spin1
2particles, while we prefer to have γ5diagonal when we deal with fast
moving spin1
2particles .
I note in passing that if we define σμ≡(I,/vectorσ)and¯σμ≡(I,−/vectorσ)we can write
γμ=/parenleftBigg0σμ
¯σμ0/parenrightBigg
more compactly in the Weyl basis. (We develop this further in appendix E.)
100 | II. Dirac and the Spinor
Chirality or handedness
Regardless of whether a Dirac field ψ(x) is massive or massless, it is enormously useful to
decompose ψinto left and right handed fields ψ(x)=ψL(x)+ψR(x)≡1
2(1−γ5)ψ(x) +
1
2(1+γ5)ψ(x). As an exercise, show that you can write the Dirac Lagrangian as
L=¯ψ(i/negationslash∂−m)ψ=¯ψLi/negationslash∂ψL+¯ψRi/negationslash∂ψR−m(¯ψLψR+¯ψRψL) (24)
The kinetic energy connects left to left and right to right, while the mass term connects
left to right and right to left.
The transformation ψ→eiθψleaves the Lagrangian Linvariant. Applying Noether’s
theorem, we obtain the conserved current associated with this symmetry Jμ=¯ψγμψ.
Projecting into left and right handed fields we see that they transform the same way:ψ
L→eiθψLandψR→eiθψR.
Ifm=0,Lenjoys an additional symmetry, known as a chiral symmetry, under which
ψ→eiφγ5ψ. Noether’s theorem tells us that the axial current J5μ≡¯ψγμγ5ψis conserved.
The left and right handed fields transform in opposite ways: ψL→e−iφψLandψR→
eiφψR. These points are particularly obvious when Lis written in terms of ψLandψR,a s
in (24).
In 1956 Lee and Yang proposed that the weak interaction does not preserve parity. It was
eventually realized (with these four words I brush over a beautiful chapter in the historyof particle physics; I urge you to read about it!) that the weak interaction Lagrangian hasthe generic form
L=G¯ψ1Lγμψ2L¯ψ3Lγμψ4L (25)
where ψ1, 2, 3, 4 denotes four Dirac fields and Gthe Fermi coupling constant. This La-
grangian clearly violates parity: Under a spatial reflection, left handed fields are trans-formed into right handed fields and vice versa.
Incidentally, henceforth when I say a Lagrangian has a certain form, I will usually
indicate only one or more of the relevant terms in the Lagrangian, as in (25). The otherterms in the Lagrangian, such as ¯ψ
1(i/negationslash∂−m1)ψ1, are understood. If the term is not
hermitean, then it is understood that we also add its hermitean conjugate.
Interactions
As we saw in (25) given the classification of bilinears in the spinor field you worked outin an exercise it is easy to introduce interactions. As another example, we can couplea scalar field ϕto the Dirac field by adding the term gϕ¯ψψ (with gsome coupling
constant) to the Lagrangian L=¯ψ(i/negationslash∂−m)ψ (and of course also adding the Lagrangian
forϕ). Similarly, we can couple a vector field A
μby adding the term eAμ¯ψγμψ.W e
note that in this case we can introduce the covariant derivative Dμ=∂μ−ieAμand write
II.1. The Dirac Equation | 101
L=¯ψ(i/negationslash∂−m)ψ+eAμ¯ψγμψ=¯ψ(iγμDμ−m)ψ . Thus, the Lagrangian for a Dirac field
interacting with a vector field of mass μreads
L=¯ψ(iγμDμ−m)ψ−1
4FμνFμν−1
2μ2AμAμ(26)
If the mass μvanishes, this is the Lagrangian for quantum electrodynamics. Varying with
respect to ¯ψ, we obtain the Dirac equation in the presence of an electromagnetic field:
[iγμ(∂μ−ieAμ)−m]ψ=0 (27)
Charge conjugation and antimatter
With coupling to the electromagnetic field, we have the concept of charge and hence of
charge conjugation. Let us try to flip the charge e. Take the complex conjugate of (27):
[−iγμ∗(∂μ+ieAμ)−m]ψ∗=0. Complex conjugating (2) we see that the −γμ∗also satisfy
the Clifford algebra and thus must be the γμmatrices expressed in a different basis, that
is, there exists a matrix Cγ0(the notation with an explicit factor of γ0is standard; see
below) such that −γμ∗=(Cγ0)−1γμ(Cγ0). Plugging in, we find that
[iγμ(∂μ+ieAμ)−m]ψc=0 (28)
where we have defined ψc≡Cγ0ψ∗. Thus, if ψis the field of the electron, then ψcis the
field of a particle with a charge opposite to that of the electron but with the same mass,namely the positron.
The discovery of antimatter was one of the most momentous in twentieth-century
physics. We will discuss antimatter in more detail in the next chapter.
It may be instructive to look at the specific form of the charge conjugation matrix C.W e
can write the defining equation for CasCγ
0γμ∗γ0C−1=−γμ. Complex conjugating the
equation (γμ)†=γ0γμγ0, we obtain (γμ)T=γ0γμ∗γ0ifγ0is real. Thus,
(γμ)T=−C−1γμC (29)
which explains why Cis defined with a γ0attached.
In both the Dirac and the Weyl bases γ2is the only imaginary gamma matrix. Then the
defining equation for Cjust says that Cγ0commutes with γ2but anticommutes with the
other three γmatrices. So evidently C=γ2γ0[up to an arbitrary phase not fixed by (29)]
and indeed γ2γμ∗γ2=γμ. Note that we have the simple (and satisfying) relation
ψc=γ2ψ∗(30)
You can easily convince yourself (exercise II.1.9) that the charge conjugate of a left
handed field is right handed and vice versa. As we will see later, this fact turns out to
be crucial in the construction of grand unified theory. Experimentally, it is known that theneutrino is left handed. Thus, we can now predict that the antineutrino is right handed.
102 | II. Dirac and the Spinor
Furthermore, ψctransforms as a spinor. Let’s check: Under a Lorentz transformation
ψ→e−(i/4 )ωμνσμνψ, complex conjugating we have ψ∗→e+(i/4 )ωμν(σμν)∗ψ∗; hence ψc→
γ2e+(i/4 )ωμν(σμν)∗ψ∗=e−(i/4 )ωμνσμνψc. [Recall from (10) that σμνis defined with an explicit
i.]
Note that CT=γ0γ2=−Cin both the Dirac and the Weyl bases.
Majorana neutrino
Since ψctransforms as a spinor, Majorana2noted that Lorentz invariance allows not only
the Dirac equation i/negationslash∂ψ=mψ but also the Majorana equation
i/negationslash∂ψ=mψc (31)
Complex conjugating this equation and multiplying by γ2, we have −γ2iγμ∗∂μψ∗=
γ2m(−γ2)ψ, that is, i/negationslash∂ψc=mψ . Thus, −∂2ψ=i/negationslash∂(i/negationslash∂ψ)=i/negationslash∂mψc=m2ψ. As we antic-
ipated, mis indeed the mass, known as a Majorana mass, of the particle associated with
ψ.
The Majorana equation (31) can be obtained from the Lagrangian3
L=¯ψi/negationslash∂ψ−1
2m(ψTCψ+¯ψC¯ψT) (32)
upon varying ¯ψ.
Sinceψandψccarry opposite charge, the Majorana equation, unlike the Dirac equation,
can only be applied to electrically neutral fields. However, as ψcis right handed if ψis left
handed, the Majorana equation, again unlike the Dirac equation, preserves handedness.Thus, the Majorana equation is almost tailor made for the neutrino.
From its conception the neutrino was assumed to be massless, but couple of years ago
experimentalists established that it has a small but nonvanishing mass. As of this writing, itis not known whether the neutrino mass is Dirac or Majorana. We will see in chapter VII.7that a Majorana mass for the neutrino arises naturally in the SO( 10)grand unified theory.
Finally, there is the possibility that ψ=ψ
c, in which case ψis known as a Majorana
spinor.
Time reversal
Finally, we come to time reversal,4which as you probably know, is much more confusing
to discuss than parity and charge conjugation. In a famous 1932 paper Wigner showed that
2Ettore Majorana had a brilliant but tragically short career. In his early thirties, he disappeared off the coast
of Sicily during a boat trip. The precise cause of his death remains a mystery. See F. Guerra and N. Robotti, Ettore
Majorana: Aspects of His Scientific and Academic Activity .
3Upon recalling that Cis antisymmetric, you may have worried that ψTCψ=Cαβψαψβvanishes. In future
chapters we will learn that ψhas to be treated as anticommuting “Grassmannian numbers.”
4Incidentally, I do not feel that we completely understand the implications of time-reversal invariance. See
A. Zee, “Night thoughts on consciousness and time reversal,” in: Art and Symmetry in Experimental Physics : pp.
246–249 .
II.1. The Dirac Equation | 103
time reversal is represented by an antiunitary operator. Since this peculiar feature already
appears in nonrelativistic quantum physics, it is in some sense not the responsibility of abook on relativistic quantum field theory to explain time reversal as an antiunitary operator.Nevertheless, let me try to be as clear as possible. I adopt the approach of “letting thephysics, namely the equations, lead us.”
Take the Schr ¨odinger equation i(∂/∂t)/Psi1(t) =H/Psi1(t) (and for definiteness, think of
H=−(1/2m)∇
2+V(/vectorx), just simple one particle nonrelativistic quantum mechanics.)
We suppress the dependence of /Psi1on/vectorx. Consider the transformation t→t/prime=−t.W e
want to find a /Psi1/prime(t/prime)such that i(∂/∂t/prime)/Psi1/prime(t/prime)=H/Psi1/prime(t/prime). Write /Psi1/prime(t/prime)=T /Psi1(t), where T
is some operator to be determined (up to some arbitrary phase factor η). Plugging in, we
havei[∂/∂(−t)]T /Psi1(t) =HT/Psi1(t) . Multiply by T−1, and we obtain T−1(−i)T (∂/∂t)/Psi1(t) =
T−1HT/Psi1(t) . Since Hdoes not involve time in any way, we want T−1H=HT−1. Then
T−1(−i)T (∂/∂t)/Psi1(t) =H/Psi1(t) . We are forced to conclude, as Wigner was, that
T−1(−i)T =i (33)
Speaking colloquially, we can say that in quantum physics time goes with an iand so
flipping time means flipping ias well.
LetT=UK , where Kcomplex conjugates everything to its right. Then T−1=KU−1
and (33) holds if U−1iU=i, that is, if U−1is just an ordinary (unitary) operator that does
nothing to i. We will determine Uas we go along. The presence of Kmakes T“antiunitary.”
We check that this works for a spinless particle in a plane wave state /Psi1(t)=ei(/vectork./vectorx−Et).
Plugging in, we have /Psi1/prime(t/prime)=T /Psi1(t) =UK/Psi1(t) =U/Psi1∗(t)=Ue−i(/vectork./vectorx−Et); since /Psi1has
only one component, Uis just a phase factor5ηthat we can choose to be 1. Rewriting, we
have/Psi1/prime(t)=e−i(/vectork./vectorx+Et)=ei(−/vectork./vectorx−Et). Indeed, /Psi1/primedescribes a plane wave moving in the
opposite direction. Crucially, /Psi1/prime(t)∝e−iEtand thus has positive energy as it should. Note
that acting on a spinless particle T2=UKUK =UU∗K2=+ 1.
Next consider a spin1
2nonrelativistic electron. Acting with Ton the spin-up state/parenleftBig1
0/parenrightBig
we want to obtain the spin-down state/parenleftBig0
1/parenrightBig
. Thus, we need a nontrivial matrix U=ησ2
to flip the spin:
T/parenleftBigg1
0/parenrightBigg
=U/parenleftBigg1
0/parenrightBigg
=iη/parenleftBigg0
1/parenrightBigg
Similarly, Tacting on the spin-down state produces the spin-up state. Note that acting on
a spin1
2particle
T2=ησ2Kησ 2K=ησ2η∗σ∗
2KK=− 1
This is the origin of Kramer’s degeneracy: In a system with an odd number of electrons
in an electric field, no matter how complicated, each energy level is twofold degenerate.The proof is very simple: Since the system is time reversal invariant, /Psi1andT/Psi1 have the
same energy. Suppose they actually represent the same state. Then T/Psi1=e
iα/Psi1, but then
5It is a phase factor rather than an arbitrary complex number because we require that |/Psi1/prime|2=|/Psi1|2.
104 | II. Dirac and the Spinor
T2/Psi1=T( T/Psi1) =Teiα/Psi1=e−iαT/Psi1=/Psi1/negationslash=−/Psi1.S o/Psi1andT/Psi1 must represent two distinct
states.
All of this is beautiful stuff, which as I noted earlier you could and should have learned
in a decent course on quantum mechanics. My responsibility here is to show you how itworks for the Dirac equation. Multiplying (1) by γ
0from the left, we have i(∂/∂t)ψ(t) =
Hψ(t) withH=−iγ0γi∂i+γ0m. Once again, we want i(∂/∂t/prime)ψ/prime(t/prime)=Hψ/prime(t/prime)with
ψ/prime(t/prime)=T ψ(t) andTsome operator to be determined. The discussion above carries
over if T−1HT=H, that is, KU−1HUK =H. Thus, we require KU−1γ0UK=γ0and
KU−1(iγ0γi)UK=iγ0γi. Multiplying by Kon the left and on the right, we see that
we have to solve for a Usuch that U−1γ0U=γ0∗andU−1γiU=−γi∗. We now restrict
ourselves to the Dirac and Weyl bases, in both of which γ2is the only imaginary guy. Okay,
what flips γ1andγ3but not γ0andγ2? Well, U=ηγ1γ3(withηan arbitrary phase factor)
works:
ψ/prime(t/prime)=ηγ1γ3Kψ(t) (34)
Since the γi’s are the same in both the Dirac and the Weyl bases, in either we have from
(4)
U=η(σ1⊗iτ2)(σ3⊗iτ2)=ηiσ2⊗1
As we expect, acting on the 2-component spinors contained in ψ, the time reversal operator
Tinvolves multiplying by iσ2. Note also that as in the nonrelativistic case T2ψ=−ψ.
It may not have escaped your notice that γ0appears in the parity operator (18), γ2in
charge conjugation (30), and γ1γ3in time reversal (34). If we change a Dirac particle to its
antiparticle and flip spacetime, γ5appears.
CPT theorem
There exists a profound theorem stating that any local Lorentz invariant field theory must
be invariant under6CPT , the combined action of charge conjugation, parity, and time
reversal. The pedestrian proof consists simply of checking that any Lorentz invariant localinteraction you can write down [such as (25)], while it may break charge conjugation,parity, or time reversal separately, respects CPT . The more fundamental proof involves
considerable formal machinery that I will not develop here. You are urged to read aboutthe phenomenological study of charge conjugation, parity, time reversal, and CPT , surely
one of the most fascinating chapters in the history of physics.
7
6A rather pedantic point, but potentially confusing to some students, is that I distinguish carefully between the
action of charge conjugation Cand the matrix C: Charge conjugation Cinvolves taking the complex conjugate of
ψand then scrambling the components with Cγ0. Similarly, I distinguish between the operation of time reversal
Tand the matrix T.
7See, e.g., J. J. Sakurai, Invariance Principles and Elementary Particles and E. D. Commins, Weak Interactions .
II.1. The Dirac Equation | 105
Two stories
I end this chapter with two of my favorite physics stories—one short and one long.
Paul Dirac was notoriously a man of few words. Dick Feynman told the story that when
he first met Dirac at a conference, Dirac said after a long silence, “I have an equation; doyou have one too?”
Enrico Fermi did not usually take notes, but during the 1948 Pocono conference (see
chapter I.7) he took voluminous notes during Julian Schwinger’s lecture. When he gotback to Chicago, he assembled a group consisting of two professors, Edward Teller andGregory Wentzel, and four graduate students, Geoff Chew, Murph Goldberger, MarshallRosenbluth, and Chen-Ning Yang (all to become major figures later). The group met inFermi’s office several times a week, a couple of hours each time, to try to figure out whatSchwinger had done. After 6 weeks, everyone was exhausted. Then someone asked, “Didn’tFeynman also speak?” The three professors, who had attended the conference, said yes.But when pressed, not Fermi, nor Teller, nor Wentzel could recall what Feynman had said.All they remembered was his strange notation: pwith a funny slash through it.
8
Exercises
II.1.1 Show that the following bilinears in the spinor field ¯ψψ ,¯ψγμψ,¯ψσμνψ,¯ψγμγ5ψ, and ¯ψγ5ψtrans-
form under the Lorentz group and parity as a scalar, a vector, a tensor, a pseudovector or axial vector, and
a pseudoscalar, respectively. [Hint: For example, ¯ψγμγ5ψ→¯ψ[1+(i/4)ωσ ]γμγ5[1−(i/4)ωσ ]ψunder
an infinitesimal Lorentz transformation and →¯ψγ0γμγ5γ0ψunder parity. Work out these transforma-
tion laws and show that they define an axial vector.]
II.1.2 Write all the bilinears in the preceding exercise in terms of ψLandψR.
II.1.3 Solve (/negationslashp−m)ψ(p) =0 explicitly (by rotational invariance it suffices to solve it for /vectorpalong the 3rd
direction, say). Verify that indeed χis much smaller than φfor a slowly moving electron. What happens
for a fast moving electron?
II.1.4 Exploiting the fact that χis much smaller than φfor a slowly moving electron, find the approximate
equation satisfied by φ.
II.1.5 For a relativistic electron moving along the z-axis, perform a rotation around the z-axis. In other words,
study the effect of e−(i/4 )ωσ12onψ(p) and verify the assertion in the text regarding ψLandψR.
II.1.6 Solve the massless Dirac equation.
II.1.7 Show explicitly that (25) violates parity.
II.1.8 The defining equation for Cevidently fixes Conly up to an overall constant. Show that this constant is
fixed by requiring (ψc)c=ψ.
8C. N. Yang, Lecture at the Schwinger Memorial Session of the American Physical Society meeting in
Washington D. C., 1995.
106 | II. Dirac and the Spinor
II.1.9 Show that the charge conjugate of a left handed field is right handed and vice versa.
II.1.10 Show that ψCψ is a Lorentz scalar.
II.1.11 Work out the Dirac equation in (1 +1)-dimensional spacetime.
II.1.12 Work out the Dirac equation in (2 +1)-dimensional spacetime. Show that the apparently innocuous
mass term violates parity and time reversal. [Hint: The three γμ’s are just the three Pauli matrices with
appropriate factors of i.]
II.2 Quantizing the Dirac Field
Anticommutation
We will use the canonical formalism of chapter I.8 to quantize the Dirac field.
Long and careful study of atomic spectroscopy revealed that the wave function of two
electrons had to be antisymmetric upon exchange of their quantum numbers. It followsthat we cannot put two electrons into the same energy level so that they will have thesame quantum numbers. In 1928 Jordan and Wigner showed how this requirement of anantisymmetric wave function can be formalized by having the creation and annihilationoperators for electrons satisfy anticommutation rather than commutation relations as in(I.8.12).
Let us start out with a state with no electron |0/angbracketrightand denote by b
†
αthe operator creating
an electron with the quantum numbers α. In other words, the state b†
α|0/angbracketrightis the state
with an electron having the quantum numbers α. Now suppose we want to have another
electron with the quantum numbers β, so we construct the state b†
βb†
α|0/angbracketright. For this to be
antisymmetric upon interchanging αandβwe must have
{b†
α,b†
β}≡b†
αb†
β+b†
βb†
α=0 (1)
Upon hermitean conjugation, we have {bα,bβ}=0. In particular, b†
αb†
α=0, so that we
cannot create two electrons with the same quantum numbers.
To this anticommutation relation we add
{bα,b†
β}=δ αβ (2)
One way of arguing for this is to say that we would like the number operator to be
N=/summationtext
αb†
αbα, just as in the bosonic case. Show with one line of algebra that [ AB,C]=
A[B,C]+[A,C]Bor [AB,C]=A{B ,C}−{A,C}B. (A heuristic way of remembering
the minus sign in the anticommuting case is that we have to move CpastBin order for
Cto do its anticommuting with A.)For the desired number operator to work we need
[/summationtext
αb†
αbα,b†
β]=+b†
β(so that as usual N|0/angbracketright=0, and Nb†
β|0/angbracketright=b†
β|0/angbracketright)and so we have (2).
108 | II. Dirac and the Spinor
The Dirac field
Let us now turn to the free Dirac Lagrangian
L=¯ψ(i/negationslash∂−m)ψ (3)
The momentum conjugate to ψisπα=δL/δ∂tψα=iψ†
α. We anticipate that the correct
canonical procedure requires imposing the anticommutation relation:
{ψα(/vectorx,t),ψ†
β(/vector0,t)}=δ(3)(/vectorx)δαβ (4)
We will derive this below.
The Dirac field satisfies
(i/negationslash∂−m)ψ=0 (5)
Plugging in plane waves u(p ,s)e−ipxandv(p ,s)eipxforψ, we have
(/negationslashp−m)u(p ,s)=0 (6)
and
(/negationslashp+m)v(p ,s)=0 (7)
The index s=± 1 reminds us that each of these two equations has two solutions, spin
up and spin down. Evidently, under a Lorentz transformation the two spinors uandv
transform in the same way as ψ. Thus, if we define ¯u≡u†γ0and¯v≡v†γ0, then ¯uuand
¯vvare Lorentz scalars.
This subject is full of “peculiar” signs and so I will proceed very carefully and show you
how every sign makes sense.
First, since (6) and (7) are linear we have to fix the normalization of uandv. Since
¯u(p ,s)u(p ,s)and¯v(p ,s)v(p ,s)are Lorentz scalars, the normalization condition we
impose on them in the rest frame will hold in any frame.
Our strategy is to do things in the rest frame using a particular basis and then invoke
Lorentz invariance and basis independence. In the rest frame, (6) and (7) reduce to
(γ0−1)u=0 and (γ0+1)v=0. In particular, in the Dirac basis γ0=/parenleftBigI 0
0−I/parenrightBig
, so the
two independent spinors u(labeled by spin s=± 1)have the form
⎛
⎜⎜⎜⎜⎜⎝1
000⎞
⎟⎟⎟⎟⎟⎠and⎛
⎜⎜⎜⎜⎜⎝0
1
00⎞
⎟⎟⎟⎟⎟⎠
II.2. Quantizing the Dirac Field | 109
while the two independent spinors vhave the form
⎛
⎜⎜⎜⎜⎜⎝0
0
1
0⎞
⎟⎟⎟⎟⎟⎠and⎛
⎜⎜⎜⎜⎜⎝0
00
1⎞
⎟⎟⎟⎟⎟⎠
The normalization conditions we have implicitly chosen are then ¯u(p ,s)u(p ,s)=1 and
¯v(p ,s)v(p ,s)=− 1. Note the minus sign thrust upon us. Clearly, we also have the orthog-
onality condition ¯uv=0 and ¯vu=0. Lorentz invariance and basis independence then tell
us that these four relations hold in general.
Furthermore, in the rest frame
/summationdisplay
suα(p,s)¯uβ(p,s)=/parenleftBiggI 0
00/parenrightBigg
αβ=1
2(γ0+1)αβ
and
/summationdisplay
svα(p,s)¯vβ(p,s)=/parenleftBigg00
0−I/parenrightBigg
αβ=1
2(γ0−1)αβ
Thus, in general
/summationdisplay
suα(p,s)¯uβ(p,s)=/parenleftbigg/negationslashp+m
2m/parenrightbigg
αβ(8)
and
/summationdisplay
svα(p,s)¯vβ(p,s)=/parenleftbigg/negationslashp−m
2m/parenrightbigg
αβ(9)
Another way of deriving (8) is to note that the left hand side i sa4b y4 matrix (it is like a
column vector multiplied by a row vector on the right) and so must be a linear combinationof the sixteen 4 by 4 matrices we listed in chapter II.1. Argue that γ
5andγμγ5are ruled out
by parity and that σμνis ruled out by Lorentz invariance and the fact that only one Lorentz
vector, namely pμ, is available. Hence the right hand side must be a linear combination of
/negationslashpandm. Fix the relative coefficient by acting with /negationslashp−mfrom the left. The normalization
is fixed by setting α=βand summing over α. Similarly for (9). In particular, setting α=β
and summing over α, we recover ¯v(p ,s)v(p ,s)=− 1.
We are now ready to promote ψ(x) to an operator. In analogy with (I.8.11) we expand
the field in plane waves1
ψα(x)=
/integraldisplayd3p
(2π)3
2(Ep/m)1
2/summationdisplay
s[b(p ,s)uα(p,s)e−ipx+d†(p,s)vα(p,s)eipx] (10)
(Here Ep=p0=+/radicalbig
/vectorp2+m2andpx=pμxμ.)The normalization factor (Ep/m)1
2is
slightly different from that in (I.8.11) for reasons we will see. Otherwise, the rationale
1The notation is standard. See e.g., J. A. Bjorken and S. D. Drell, Relativistic Quantum Mechanics .
110 | II. Dirac and the Spinor
for (10) is essentially the same as in (I.8.11). We integrate over momentum /vectorp, sum over
spins, expand in plane waves, and give names to the coefficients in the expansion. Because
ψis complex, we have, similar to the complex scalar field in chapter I.8, a boperator and
ad†operator.
Just as in chapter I.8, the operators bandd†must carry the same charge. Thus, if b
annihilates an electron with charge e=− |e|,d†must remove charge e; that is, it creates
a positron with charge −e=|e|.
A word on notation: in (10) b(p ,s),d†(p,s),u(p ,s), andv(p ,s)are written as functions
of the 4-momentum pbut strictly speaking they are functions of /vectorponly, with p0always
understood to be +/radicalbig
/vectorp2+m2.
Thus let b†(p,s)andb(p ,s)be the creation and annihilation operators for an electron
of momentum pand spin s. Our introductory discussion indicates that we should impose
{b(p ,s),b†(p/prime,s/prime)}=δ(3)(/vectorp−/vectorp/prime)δss/prime (11)
{b(p ,s),b(p/prime,s/prime)}=0 (12)
{b†(p,s),b†(p/prime,s/prime)}=0 (13)
There is a corresponding set of relations for d†(p,s)andd(p ,s)the creation and annihi-
lation operators for a positron, for instance,
{d(p ,s),d†(p/prime,s/prime)}=δ(3)(/vectorp−/vectorp/prime)δss/prime (14)
We now have to show that we indeed obtain (4). Write
¯ψ(0)=/integraldisplayd3p/prime
(2π)3
2(Ep/prime/m)1
2/summationdisplay
s/prime[b†(p/prime,s/prime)¯u(p/prime,s/prime)+d(p/prime,s/prime)¯v(p/prime,s/prime)]
Nothing to do but to plow ahead:
{ψ(/vectorx,0),¯ψ(0)}
=/integraldisplayd3p
(2π)3(Ep/m)/summationdisplay
s[u(p ,s)¯u(p ,s)ei/vectorp./vectorx+v(p ,s)¯v(p ,s)e−i/vectorp./vectorx]
if we take bandb†to anticommute with dandd†. Using (8) and (9) we obtain
{ψ(/vectorx,0),¯ψ(0)}=/integraldisplayd3p
(2π)3(2Ep)[(/negationslashp+m)ei/vectorp./vectorx+(/negationslashp−m)e−i/vectorp./vectorx]
=/integraldisplayd3p
(2π)3(2Ep)2p0γ0e−i/vectorp./vectorx=γ0δ(3)(/vectorx)
which is just (4) slightly disguised.
Similarly, writing schematically, we have {ψ,ψ}=0 and {ψ†,ψ†}=0.
We are of course free to normalize the spinors uandvhowever we like. One alternative
normalization is to define uandvas the uandvgiven here multiplied by (2m)1
2, thus
changing (8) and (9) to/summationtext
su¯u=/negationslashp+mand/summationtext
sv¯v=/negationslashp−m. Multiplying the numerator and
denominator in (10) by (2m)1
2, we see that the normalization factor (Ep/m)1
2is changed to
(2Ep)1
2[thus making it the same as the normalization factor for the scalar field in (I.8.11)].
II.2. Quantizing the Dirac Field | 111
This alternative normalization (let us call it “any mass normalization”) is particularly
convenient when we deal with massless spin-1
2particles: we could set m=0 everywhere
without ever encountering min the denominator as in (8) and (9).
The advantage of the normalization used here (let us call it “rest normalization”) is that
the spinors assume simple forms in the rest frame, as we have just seen. This would proveto be advantageous when we calculate the magnetic moment of the electron in chapter
III.6, for example. Of course, multiplying and dividing here and there by (2m)
1
2is a trivial
operation, and there is not much sense in arguing over the relative advantages of onenormalization over another.
In chapter II.6 we will calculate electron scattering at energies high compared to the
massm, so that effectively we could set m=0. Actually, even then, “rest normalization”
has the slight advantage of providing a (rather weak) check on the calculation. We set m=0
everywhere we can, such as in the numerator of (8) and (9), but not where we can’t, such asin the denominator. Then mmust cancel out in physical quantities such as the differential
scattering cross section.
Energy of the vacuum
An important exercise at this point is to calculate the Hamiltonian starting with theHamiltonian density
H=π∂ψ
∂t−L=¯ψ(i/vectorγ./vector∂+m)ψ (15)
Inserting (10) into this expression and integrating, we have
H=/integraldisplay
d3xH=/integraldisplay
d3x¯ψ(i/vectorγ./vector∂+m)ψ=/integraldisplay
d3x¯ψiγ0∂ψ
∂t(16)
which works out to be
H=/integraldisplay
d3p/summationdisplay
sEp[b†(p,s)b(p ,s)−d(p ,s)d†(p,s)] (17)
We can see the all important minus sign in (17) schematically: In (16) ¯ψgives a factor
∼(b†+d), while ∂/∂t acting on ψbrings down a relative minus sign giving ∼(b−d†),
thus giving us ∼(b†+d)(b−d†)∼b†b−dd†(orthogonality between spinors vu=0 kills
the cross terms).
To bring the second term in (17) into the right order, we anticommute −d(p ,s)d†(p,s)
=d†(p,s)d(p ,s)−δ(3)(/vector0)so that
H=/integraldisplay
d3p/summationdisplay
sEp[b†(p,s)b(p ,s)+d†(p,s)d(p ,s)]
−δ(3)(/vector0)/integraldisplay
d3p/summationdisplay
sEp (18)
The first two terms tell us that each electron and each positron of momentum pand spin
shas exactly the same energy Ep, as it should. But what about the last term? That δ(3)(/vector0)
should fill us with fear and loathing.
112 | II. Dirac and the Spinor
It is OK: Noting that δ(3)(/vectorp)=[1/(2π)3]/integraltext
d3xei/vectorp/vectorx, we see that δ(3)(/vector0)=[1/(2π)3]/integraltext
d3x
(we encounter the same maneuver in exercise I.8.2) and so the last term contributes to H
E0=−1
h3/integraldisplay
d3x/integraldisplay
d3p/summationdisplay
s2(1
2Ep) (19)
(since in natural units we have /planckover2pi=1 and hence h=2π). We have an energy −1
2Epin each
unit-size phase-space cell (1/h3)d3xd3pin the sense of statistical mechanics, for each spin
and for the electron and positron separately (hence the factor of 2) . This infinite additivetermE
0is precisely the analog of the zero point energy1
2/planckover2piωof the harmonic oscillator you
encountered in your quantum mechanics course. But it comes in with a minus sign!
The sign is bizarre and peculiar! Each mode of the Dirac field contributes −1
2/planckover2piωto
the vacuum energy. In contrast, each mode of a scalar field contributes1
2/planckover2piωas we saw
in chapter I.8. This fact is of crucial importance in the development of supersymmetry,which we will discuss in chapter VIII.4.
Fermion propagating through spacetime
In analogy with (I.8.14), the propagator for the electron is given by iSαβ(x)≡
/angbracketleft0|Tψα(x)¯ψβ(0)|0/angbracketright, where the argument of ¯ψhas been set to 0 by translation invariance.
As we will see, the anticommuting character of ψrequires us to define the time-ordered
product with a minus sign, namely
T ψ(x) ¯ψ(0)≡θ(x0)ψ(x) ¯ψ(0)−θ(−x0)¯ψ(0)ψ(x) (20)
Referring to (10), we obtain for x0>0,
iS(x)=/angbracketleft0|ψ(x)¯ψ(0)|0/angbracketright=/integraldisplayd3p
(2π)3(Ep/m)/summationdisplay
su(p ,s)¯u(p ,s)e−ipx
=/integraldisplayd3p
(2π)3(Ep/m)/negationslashp+m
2me−ipx
Forx0<0, we have to be a bit careful about the spinorial indices:
iSαβ(x)=− /angbracketleft 0|¯ψβ(0)ψα(x)|0/angbracketright
=−/integraldisplayd3p
(2π)3(Ep/m)/summationdisplay
s¯vβ(p,s)vα(p,s)e−ipx
=−/integraldisplayd3p
(2π)3(Ep/m)(/negationslashp−m
2m)αβe−ipx
using the identity (9).
Putting things together we obtain
iS(x)=/integraldisplayd3p
(2π)3(Ep/m)/bracketleftbigg
θ(x0)/negationslashp+m
2me−ipx−θ(−x0)/negationslashp−m
2me+ipx/bracketrightbigg
(21)
We will now show that this fermion propagator can be written more elegantly as a
4-dimensional integral:
II.2. Quantizing the Dirac Field | 113
iS(x)=i/integraldisplayd4p
(2π)4e−ip.x/negationslashp+m
p2−m2+iε=/integraldisplayd4p
(2π)4e−ip.x i
/negationslashp−m+iε(22)
To show that (22) is indeed equivalent to (21) we go through essentially the same steps as
after (I.8.14). In the complex p0plane the integrand has poles at p0=±/radicalbig
/vectorp2+m2−iε/similarequal
±(Ep−iε).F o rx0>0 the factor e−ip0x0tells us to close the contour in the lower half-plane.
We go around the pole at +(Ep−iε)clockwise and obtain
iS(x)=(−i)i/integraldisplayd3p
(2π)3e−ip.x/negationslashp+m
2Ep
producing the first term in (21). For x0<0 we are now told to close the contour in the
upper half-plane and thus we go around the pole at −(Ep−iε)anticlockwise. We obtain
iS(x)=i2/integraldisplayd3p
(2π)3e+iEpx0+i/vectorp./vectorx 1
−2Ep(−Epγ0−/vectorp/vectorγ+m)
and flipping /vectorpwe have
iS(x)=−/integraldisplayd3p
(2π)3eip.x1
2Ep(Epγ0−/vectorp/vectorγ−m)=−/integraldisplayd3p
(2π)3eip.x1
2Ep(/negationslashp−m)
precisely the second term in (21) with the minus sign and all. Thus, we must define the
time-ordered product with the minus sign as in (20).
After all these steps, we see that in momentum space the fermion propagator has the
elegant form
iS(p) =i
/negationslashp−m+iε(23)
This makes perfect sense: S(p) comes out to be the inverse of the Dirac operator /negationslashp−m,
just as the scalar boson propagator D(k)=1/(k2−m2+iε) is the inverse of the Klein-
Gordon operator k2−m2.
Poetic but confusing metaphors
In closing this chapter let me ask you some rhetorical questions. Did I speak of an
electron going backward in time? Did I mumble something about a sea of negative energyelectrons? This metaphorical language, when used by brilliant minds, the likes of Diracand Feynman, was evocative and inspirational, but unfortunately confused generationsof physics students and physicists. The presentation given here is in the modern spirit,which seeks to avoid these potentially confusing metaphors.
Exercises
II.2.1 Use Noether’s theorem to derive the conserved current Jμ=¯ψγμψ. Calculate [Q ,ψ], thus showing that
bandd†must carry the same charge.
II.2.2 Quantize the Dirac field in a box of volume of Vand show that the vacuum energy E0is indeed
proportional to V. [Hint: The integral over momentum/integraltext
d3pis replaced by a sum over discrete values
of the momentum.]
II.3 Lorentz Group and Weyl Spinors
The Lorentz algebra
In chapter II.1 we followed Dirac’s brilliantly idiosyncratic way of deriving his equation.
We develop here a more logical and mathematical theory of the Dirac spinor. A deeperunderstanding of the Dirac spinor not only gives us a certain satisfaction, but is alsoindispensable, as we will see later, in studying supersymmetry, one of the foundationalconcepts of superstring theory; and of course, most of the fundamental particles such asthe electron and the quarks carry spin
1
2and are described by spinor fields.
Let us begin by reminding ourselves how the rotation group works. The three generators
Ji(i=1, 2, 3 or x,y,z) of the rotation group satisfy the commutation relation
[Ji,Jj]=i/epsilon1ijkJk (1)
When acting on the spacetime coordinates, written as a column vector
⎛
⎜⎜⎜⎜⎜⎝x
0
x1
x2
x3⎞
⎟⎟⎟⎟⎟⎠
the generators of rotations are represented by the hermitean matrices
J1=⎛
⎜⎜⎜⎜⎜⎝000 0
000 0000−i00 i 0⎞
⎟⎟⎟⎟⎟⎠(2)
withJ2andJ3obtained by cyclic permutations. You should verify by laboriously multi-
plying these three matrices that (1) is satisfied. Note that the signs of Jiare fixed by the
commutation relation (1).
II.3. Lorentz Group and Spinors | 115
Now add the Lorentz boosts. A boost in the x≡x1direction transforms the spacetime
coordinates:
t/prime=(cosh ϕ) t+(sinh ϕ) x ;x/prime=(sinh ϕ) t+(cosh ϕ) x (3)
or for infinitesimal ϕ
t/prime=t+ϕx;x/prime=x+ϕt (4)
In other words, the infinitesimal generator of a Lorentz boost in the xdirection is repre-
sented by the hermitean matrix (x0≡tas usual)
iK1=⎛
⎜⎜⎜⎜⎜⎝0100
1000
00000000⎞
⎟⎟⎟⎟⎟⎠(5)
Similarly,
iK 2=⎛
⎜⎜⎜⎜⎜⎝0010
0000
1000
0000⎞
⎟⎟⎟⎟⎟⎠(6)
I leave it to you to write down K3. Note that Kiis defined to be antihermitean.
Check that [ Ji,Kj]=i/epsilon1ijkKk. To see that this implies that the boost generators Ki
transform as a 3-vector /vectorKunder rotation, as you would expect, apply a rotation through
an infinitesimal angle θaround the 3-axis. Then (you might wish to review the material in
appendix B at this point) K1→eiθJ 3K1e−iθJ 3=K1+iθ[J3,K1]+O(θ2)=K1+iθ(iK2)+
O(θ2)=cosθK1−sinθK2to the order indicated.
You are now about to do one of the most significant calculations in the history of
twentieth century physics. By brute force compute [ K1,K2], evidently an antisymmetric
matrix. You will discover that it is equal to −iJ 3. Two Lorentz boosts produce a rotation!
(You might recall from your course on electromagnetism that this mathematical fact isresponsible for the physics of the Thomas precession.)
Mathematically, the generators of the Lorentz group satisfy the following algebra [known
to the cognoscenti as SO( 3, 1)]:
[Ji,Jj]=i/epsilon1ijkJk (7)
[Ji,Kj]=i/epsilon1ijkKk (8)
[Ki,Kj]=−i/epsilon1ijkJk (9)
Note the all-important minus sign!
116 | II. Dirac and the Spinor
How do we study this algebra? The crucial observation is that the algebra falls apart into
two pieces if we form the combinations J±i≡1
2(Ji±iKi). You should check that
[J+i,J+j]=i/epsilon1ijkJ+k (10)
[J−i,J−j]=i/epsilon1ijkJ−k (11)
and most remarkably
[J+i,J−j]=0 (12)
This last commutation relation tells us that J+andJ−form two separate SU( 2)algebras.
(For more, see appendix B.)
From algebra to representation
This means that you can simply use what you have already learned about angular mo-mentum in elementary quantum mechanics and the representation of SU( 2)to deter-
mine all the representations of SO( 3, 1). As you know, the representations of SU( 2)
are labeled by j=0,
1
2,1 ,3
2,.... We can think of each representation as consisting of
(2j+1)objects ψmwithm=−j,−j+1,...,j−1,jthat transform into each other
under SU( 2). It follows immediately that the representations of SO( 3, 1)are labeled by
(j+,j−)withj+andj−each taking on the values 0,1
2,1 ,3
2, . . . . Each representation
consists of (2j++1)(2j−+1)objects ψm+m−withm+=−j+,−j++1 ,..., j+−1,j+
andm−=−j−,−j−+1 ,..., j−−1,j−.
Thus, the representations of SO( 3, 1)are(0, 0),(1
2,0),(0,1
2),(1, 0),(0, 1),(1
2,1
2), and
so on, in order of increasing dimension. We recognize the 1-dimensional representation(0, 0)as clearly the trivial one, the Lorentz scalar. By counting dimensions, we expect
that the 4-dimensional representation (
1
2,1
2)has to be the Lorentz vector, the defining
representation of the Lorentz group (see exercise II.3.1).
Spinor representations
What about the representation (1
2,0)? Let us write the two objects as ψαwithα=1, 2.
Well, what does the notation (1
2,0)mean? It says that J+i=1
2(Ji+iKi)acting on ψαis
represented by1
2σiwhile J−i=1
2(Ji−iKi)acting on ψαis represented by 0. By adding
and subtracting we find that
Ji=1
2σi (13)
and
iKi=1
2σi (14)
II.3. Lorentz Group and Spinors | 117
where the equal sign means “represented by” in this context. (By convention we do not
distinguish between upper and lower indices on the 3-dimensional quantities Ji,Ki, and
σi.)Note again that Kiis anti-hermitean.
Similarly, let us denote the two objects in (0,1
2)by the peculiar symbol ¯χ˙α. I should em-
phasize the trivial but potentially confusing point that unlike the bar used in chapter II.1,the bar on ¯χ
˙αis a typographical element: Think of the symbol ¯χas a letter in the Hittite
alphabet if you like. Similarly, the symbol ˙αbears no relation to α: we do not obtain ˙α
by operating on αin any way. The rather strange notation is known informally as “dotted
and undotted” and more formally as the van der Waerden notation—a bit excessive for ourrather modest purposes at this point but I introduce it because it is the notation used insupersymmetric physics and superstring theory. (Incidentally, Dirac allegedly said that hewished he had invented the dotted and undotted notation.) Repeating the same steps asabove, you will find that on the representation (0,
1
2)we have Ji=1
2σiandiKi=−1
2σi.
The minus sign is crucial.
The 2-component spinors ψαand¯χ˙αare called Weyl spinors and furnish perfectly good
representations of the Lorentz group. Why then does the Dirac spinor have 4 components?
The reason is parity. Under parity, /vectorx→−/vectorxand/vectorp→−/vectorp, and thus /vectorJ→/vectorJand/vectorK→−/vectorK,
and so /vectorJ+↔/vectorJ−. In other words, under parity the representations (1
2,0)↔(0,1
2). There-
fore, to describe the electron we must use both of these 2-dimensional representations, orin mathematical notation, the 4-dimensional reducible representation (
1
2,0)⊕(0,1
2).
We thus stack two Weyl spinors together to form a Dirac spinor
/Psi1=/parenleftBiggψα
¯χ˙α/parenrightBigg
(15)
The spinor /Psi1(p) is of course a function of 4-momentum p[and by implication also ψα(p)
and¯χ˙α(p)] but we will suppress the pdependence for the time being. Referring to (13)
and (14) we see that acting on /Psi1the generators of rotation
/vectorJ=/parenleftBigg1
2/vectorσ 0
01
2/vectorσ/parenrightBigg
where once again the equality means “represented by,” and the generators of boost
i/vectorK=/parenleftBigg1
2/vectorσ 0
0−1
2/vectorσ/parenrightBigg
Note once again the all-important minus sign.
Parity forces us to have a 4-component spinor but we know on the other hand that the
electron has only two physical degrees of freedom. Let us go to the rest frame. We mustproject out two of the components contained in /Psi1(p
r)with the rest momentum pr≡
(m,/vector0). With the benefit of hindsight, we write the projection operator as P=1
2(1−γ0).
You are probably guessing from the notation that γ0will turn out to be one of the gamma
matrices, but at this point, logically γ0is just some 4 by 4 matrix. The condition P2=P
implies that (γ0)2=1 so that the eigenvalues of γ0are±1. Since ψα↔¯χ˙αunder parity we
naturally guess that ψαand¯χ˙αcorrespond to the left and right handed fields of chapter II.1.
118 | II. Dirac and the Spinor
We cannot simply use the projection to set for example ¯χ˙αto 0. Parity means that we should
treatψαand¯χ˙αon the same footing. We choose
γ0=/parenleftBigg01
10/parenrightBigg
,
or more explicitly
⎛
⎜⎜⎜⎜⎜⎝0010
0001
1000
0100⎞
⎟⎟⎟⎟⎟⎠
(Different choices of γ0correspond to the different basis choices discussed in chapter II.1.)
In other words, in the rest frame ψα−¯χ˙α=0. The projection to two degrees of freedom
can be written as
(γ0−1)/Psi1(p r)=0 (16)
Indeed, we recognize this as just the Weyl basis introduced in chapter II.1.
The Dirac equation
We have derived the Dirac equation, a bit in disguise!
Since our derivation is based on a step-by-step study of the spinor representation of the
Lorentz group, we know how to obtain the equation satisfied by /Psi1(p) for any p: We simply
boost. Writing /Psi1(p)=e−i/vectorϕ/vectorK/Psi1(pr), we have (e−i/vectorϕ/vectorKγ0ei/vectorϕ/vectorK−1)/Psi1(p) =0. Introducing the
notation γμpμ/m≡e−i/vectorϕ/vectorKγ0ei/vectorϕ/vectorK, we obtain the Dirac equation
(γμpμ−m)/Psi1(p) =0 (17)
You can work out the details as an exercise.
The derivation here represents the deep group theoretic way of looking at the Dirac
equation: It is a projection boosted into an arbitrary frame.
Note that this is an example of the power of symmetry, which pervades modern physics
and this book: Our knowledge of how the electron field transforms under the rotationgroup, namely that it has spin
1
2, allows us to know how it transforms under the Lorentz
group. Symmetry rules!
In appendix E we will develop the dotted and undotted notation further for later use in
the chapter on supersymmetry.
In light of your deeper group theoretic understanding it is a good idea to reread chap-
ter II.1 and compare it with this chapter.
II.3. Lorentz Group and Spinors | 119
Exercises
II.3.1 Show by explicit computation that (1
2,1
2)is indeed the Lorentz vector.
II.3.2 Work out how the six objects contained in the (1, 0)and(0, 1)transform under the Lorentz group.
Recall from your course on electromagnetism how the electric and magnetic fields /vectorEand/vectorBtransform.
Conclude that the electromagnetic field in fact transforms as (1, 0)⊕(0, 1). Show that it is parity that
once again forces us to use a reducible representation.
II.3.3 Show that
e−i/vectorϕ/vectorKγ0ei/vectorϕ/vectorK=/parenleftBigg0e−/vectorϕ/vectorσ
e/vectorϕ/vectorσ0/parenrightBigg
and
e/vectorϕ/vectorσ=coshϕ+/vectorσ.ˆϕsinhϕ
with the unit vector ˆϕ≡/vectorϕ/ϕ . Identifying /vectorp=mˆϕsinhϕ, derive the Dirac equation. Show that
γi=/parenleftBigg0σi
−σi0/parenrightBigg
II.3.4 Show that a spin3
2particle can be described by a vector-spinor /Psi1αμ, namely a Dirac spinor carrying a
Lorentz index. Find the corresponding equations of motion, known as the Rarita-Schwinger equations.[Hint: The object /Psi1
αμhas 16 components, which we need to cut down to 2 .3
2+1=4 components.]
II.4 Spin-Statistics Connection
There is no one fact in the physical world which has
a greater impact on the way things are, than the Pauli
exclusion principle.1
Degrees of intellectual incompleteness
In a course on nonrelativistic quantum mechanics you learned about the Pauli exclusion
principle2and its later generalization stating that particles with half integer spins, such as
electrons, obey Fermi-Dirac statistics and want to stay apart, while in contrast particles withinteger spins, such as photons or pairs of electrons, obey Bose-Einstein statistics and loveto stick together. From the microscopic structure of atoms to the macroscopic structureof neutron stars, a dazzling wealth of physical phenomena would be incomprehensiblewithout this spin-statistics rule. Many elements of condensed matter physics, for instance,band structure, Fermi liquid theory, superfluidity, superconductivity, quantum Hall effect,and so on and so forth, are consequences of this rule.
Quantum statistics, one of the most subtle concepts in physics, rests on the fact that in
the quantum world, all elementary particles and hence all atoms, are absolutely identicalto, and thus indistinguishable from, one other.
3It should be recognized as a triumph of
quantum field theory that it is able to explain absolute identity and indistinguishabilityeasily and naturally. Every electron in the universe is an excitation in one and the sameelectron field ψ. Otherwise, one might be able to imagine that the electrons we now
1I. Duck and E. C. G. Sudarshan, Pauli and the Spin-Statistics Theorem ,p .2 1 .
2While a student in Cambridge, E. C. Stoner came to within a hair of stating the exclusion principle.
Pauli himself in his famous paper ( Zeit. f. Physik 31: 765, 1925) only claimed to “summarize and generalize
Stoner’s idea.” However, later in his Nobel Prize lecture Pauli was characteristically ungenerous toward Stoner’scontribution. A detailed and fascinating history of the spin and statistics connection may be found in Duck and
Sudarshan, op. cit.
3Early in life, I read in one of George Gamow’s popular physics books that he could not explain quantum
statistics—all he could manage for Fermi statistics was an analogy, invoking Greta Garbo’s famous remark “Ivont to be alone.”—and that one would have to go to school to learn about it. Perhaps this spurs me, later in life,
to write popular physics books also. See A. Zee, Einstein ’s Universe ,p .x .
II.4. Spin-Statistics Connection | 121
know came off an assembly line somewhere in the early universe and could all be slightly
different owing to some negligence in the manufacturing process.
While the spin-statistics rule has such a profound impact in quantum mechanics, its
explanation had to wait for the development of relativistic quantum field theory. Imaginea civilization that for some reason developed quantum mechanics but has yet to discoverspecial relativity. Physicists in this civilization eventually realize that they have to inventsome rule to account for the phenomena mentioned above, none of which involves motionfast compared to the speed of light. Physics would have been intellectually unsatisfyingand incomplete.
One interesting criterion in comparing different areas of physics is their degree of
intellectual incompleteness.
Certainly, in physics we often accept a rule that cannot be explained until we move to the
next level. For instance, in much of physics, we take as a given the fact that the charge of theproton and the charge of electron are exactly equal and opposite. Quantum electrodynamicsby itself is not capable of explaining this striking fact either. This fact, charge quantization,can only be deduced by embedding quantum electrodynamics into a larger structure, suchas a grand unified theory, as we will see in chapter VII.6. (In chapter IV .4 we will learn thatthe existence of magnetic monopoles implies charge quantization, but monopoles do notexist in pure quantum electrodynamics.)
Thus, the explanation of the spin-statistics connection, by Fierz and by Pauli in the late
1930s, and by L ¨uders and Zumino and by Burgoyne in the late 1950s, ranks as one of the
great triumphs of relativistic quantum field theory. I do not have the space to give a generaland rigorous proof
4here. I will merely sketch what goes terribly wrong if we violate the
spin-statistics connection.
The price of perversity
A basic quantum principle states that if two observables commute then they are simul-taneously diagonalizable and hence observable. A basic relativistic principle states that iftwo spacetime points are spacelike with respect to each other then no signal can propagatebetween them, and hence the measurement of an observable at one of the points cannotinfluence the measurement of another observable at the other point.
Consider the charge density J
0=i(ϕ†∂0ϕ−∂0ϕ†ϕ)in a charged scalar field theory.
According to the two fundamental principles just enunciated, J0(/vectorx,t=0)andJ0(/vectory,t=0)
should commute for /vectorx/negationslash=/vectory. In calculating the commutator of J0(/vectorx,t=0)withJ0(/vectory,t=0),
we simply use the fact that ϕ(/vectorx,t=0)and∂0ϕ(/vectorx,t=0)commute with ϕ(/vectory,t=0)and
∂0ϕ(/vectory,t=0), so we just move the field at /vectorxsteadily past the field at /vectory. The commutator
vanishes almost trivially.
4See I. Duck and E. C. G. Sudarshan, Pauli and the Spin-Statistics Theorem , and R. F. Streater and A. S.
Wightman, PCT , Spin Statistics, and All That .
122 | II. Dirac and the Spinor
Now suppose we are perverse and quantize the creation and annihilation operators in
the expansion (I.8.11)
ϕ(/vectorx,t=0)=/integraldisplaydDk/radicalbig
(2π)D2ωk[a(/vectork)ei/vectork./vectorx+a†(/vectork)e−i/vectork./vectorx] (1)
according to anticommutation rules
{a(/vectork),a†(/vectorq)}=δ(D)(/vectork−/vectorq)
and
{a(/vectork),a(/vectorq)}=0={a†(/vectork),a†(/vectorq)}
instead of the correct commutation rules.
What is the price of perversity?Now when we try to move J
0(/vectory,t=0)pastJ0(/vectorx,t=0), we have to move the field at /vectory
past the field at /vectorxusing the anticommutator
{ϕ(/vectorx,t=0),ϕ(/vectory,t=0)}
=/integraldisplay/integraldisplaydDk/radicalbig
(2π)D2ωkdDq/radicalBig
(2π)D2ωq{[a(/vectork)ei/vectork./vectorx+a†(/vectork)e−i/vectork./vectorx], [a(/vectorq)ei/vectorq./vectory+a†(/vectorq)e−i/vectorq./vectory]}
=/integraldisplaydDk
(2π)D2ωk(ei/vectork.(/vectorx−/vectory)+e−i/vectork.(/vectorx−/vectory)) (2)
You see the problem? In a normal scalar-field theory that obeys the spin-statistics
connection, we would have computed the commutator, and then in the last expression
in (2) we would have gotten (ei/vectork.(/vectorx−/vectory)−e−i/vectork.(/vectorx−/vectory))instead of (ei/vectork.(/vectorx−/vectory)+e−i/vectork.(/vectorx−/vectory)). The
integral
/integraldisplaydDk
(2π)D2ωk(ei/vectork.(/vectorx−/vectory)−e−i/vectork.(/vectorx−/vectory))
would obviously vanish and all would be well. With the plus sign, we get in (2) a nonvan-
ishing piece of junk. A disaster if we quantize the scalar field as anticommuting! A spin 0field has to be commuting. Thus, relativity and quantum physics join hands to force thespin-statistics connection.
It is sometimes said that because of electromagnetism you do not sink through the floor
and because of gravity you do not float to the ceiling, and you would be sinking or floatingin total darkness were it not for the weak interaction, which regulates stellar burning.Without the spin-statistics connection, electrons would not obey Pauli exclusion. Matterwould just collapse.
5
Exercise
II.4.1 Show that we would also get into trouble if we quantize the Dirac field with commutation instead of
anticommutation rules. Calculate the commutator [ J0(/vectorx,0),J0(0)].
5The proof of the stability of matter, given by Dyson and Lenard, depends crucially on Pauli exclusion.
II.5Vacuum Energy, Grassmann Integrals,
and Feynman Diagrams for Fermions
The vacuum is a boiling sea of nothingness, full of sound
and fury, signifying a great deal.
—Anonymous
Fermions are weird
I developed the quantum field theory of a scalar field ϕ(x) first in the path integral
formalism and then in the canonical formalism. In contrast, I have thus far developedthe quantum field theory of the free spin
1
2fieldψ(x) only in the canonical formalism.
We learned that the spin-statistics connection forces the field operator ψ(x) to satisfy
anticommutation relations. This immediately suggests something of a mystery in writingdown the path integral for the spinor field ψ. In the path integral formalism ψ(x) is not an
operator but merely an integration variable. How do we express the fact that its operatorcounterpart in the canonical formalism anticommutes?
We presumably cannot represent ψas a commuting variable, as we did ϕ. Indeed, we will
discover that in the path integral formalism ψis to be treated not as an ordinary complex
number but as a novel kind of mathematical entity known as a Grassmann number.
If you thought about it, you would realize that some novel mathematical structure is
needed. In chapter I.3 we promoted the coordinates of point particles q
i(t)in quantum
mechanics to the notion of a scalar field ϕ(/vectorx,t). But you already know from quantum
mechanics that a spin1
2particle has the peculiar property that its wave function turns into
minus itself when rotated through 2 π. Unlike particle coordinates, half integral spin is
not an intuitive concept.
Vacuum energy
To motivate the introduction of Grassmann-valued fields I will discuss the notion of
vacuum energy. The reason for this apparently strange strategy will become clear shortly.
Quantum field theory was first developed to describe the scattering of photons and
electrons, and later the scattering of particles. Recall that in chapter I.7 while studying the
124 | II. Dirac and the Spinor
scattering of particles we encountered diagrams describing vacuum fluctuations, which
we simply neglected (see fig. I.7.12). Quite naturally, particle physicists considered thesefluctuations to be of no importance. Experimentally, we scatter particles off each other. Whocares about fluctuations in the vacuum somewhere else? It was only in the early 1970s thatphysicists fully appreciated the importance of vacuum fluctuations. We will come back tothe importance of the vacuum
1in a later chapter.
In chapters I.8 and II.2 we calculated the vacuum energy of a free scalar field and of a free
spinor field using the canonical formalism. To motivate the use of Grassmann numbersto formulate the path integral for the spinor field I will adopt the following strategy. First,I use the path integral formalism to obtain the result we already have for the free scalarfield using the canonical formalism. Then we will see that in order to produce the resultwe already have for the free spinor field we must modify the path integral.
By definition, vacuum fluctuations occur even when there are no sources to produce
particles. Thus, let us consider the generating functional of a free scalar field theory in theabsence of sources:
2
Z=/integraldisplay
Dϕei/integraltext
d4x1
2[(∂ϕ)2−m2ϕ2]=C/parenleftbigg1
det[∂2+m2]/parenrightbigg1
2
=Ce−1
2Tr log(∂2+m2)(1)
For the first equality we used (I.2.15) and absorbed inessential factors into the constant C.
In the second equality we used the important identity
detM=eTr log M(2)
which you encountered in exercise I.11.2.
Recall that Z=/angbracketleft0|e−iHT|0/angbracketright(with T→∞ understood so that we integrate over all
of spacetime in (1)), which in this case is just e−iETwithEthe energy of the vacuum.
Evaluating the trace in (1)
TrO=/integraldisplay
d4x/angbracketleftx|O|x/angbracketright
=/integraldisplay
d4x/integraldisplayd4k
(2π)4/integraldisplayd4q
(2π)4/angbracketleftx|k/angbracketright/angbracketleftk|O|q/angbracketright/angbracketleftq|x/angbracketright
we obtain
iET=1
2VT/integraldisplayd4k
(2π)4log(k2−m2+iε)+A
where Ais an infinite constant corresponding to the multiplicative factor Cin (1). Recall
that in the derivation of the path integral we had lots of divergent multiplicative factors;this is where they can come in. The presence of Ais a good thing here since it solves a
problem you might have noticed: The argument of the log is not dimensionless. Let usdefine m
/primeby writing
1Indeed, we have already discussed one way to observe the effects of vacuum fluctuations in chapter I.9.
2Strictly speaking, to render the expressions here well defined we should replace m2bym2−iεas discussed
earlier.
II.5. Feynman Diagrams for Fermions | 125
A=−1
2VT/integraldisplayd4k
(2π)4log(k2−m/prime2+iε)
In other words, we do not calculate the vacuum energy as such, but only the difference
between it and the vacuum energy we would have had if the particle had mass m/primeinstead
ofm. The arbitrarily long time Tcancels out and Eis proportional to the volume of space
V, as might be expected. Thus, the (difference in) vacuum energy density is
E
V=−i
2/integraldisplayd4k
(2π)4log/bracketleftbiggk2−m2+iε
k2−m/prime2+iε/bracketrightbigg
(3)
=−i
2/integraldisplayd3k
(2π)3/integraldisplaydω
2πlog/bracketleftBigg
ω2−ω2
k+iε
ω2−ω/prime2
k+iε/bracketrightBigg
where ω/prime
k≡+/radicalbig
/vectork2+m/prime2. We treat the (convergent) integral over ωby integrating by parts:
/integraldisplaydω
2πdω
dωlog/bracketleftBigg
ω2−ω2
k+iε
ω2−ω/prime2
k+iε/bracketrightBigg
=− 2/integraldisplaydω
2πω/bracketleftBigg
ω
ω2−ω2
k+iε−(ωk→ω/prime
k)/bracketrightBigg
=−i2ω2
k(1
−2ωk)−(ωk→ω/prime
k)
=+i(ωk−ω/prime
k) (4)
Indeed, restoring /planckover2piwe get the result we want:
E
V=/integraldisplayd3k
(2π)3(1
2/planckover2piωk−1
2/planckover2piω/prime
k) (5)
We had to go through a few arithmetical steps to obtain this result, but the important point
is that using the path integral formalism we have managed to obtain a result previouslyobtained using the canonical formalism.
A peculiar sign for fermions
Our goal is to figure out the path integral for the spinor field. Recall from chapter II.2 that
the vacuum energy of the spinor field comes out to have the opposite sign to the vacuum
energy of the scalar field, a sign that surely ranks among the “top ten” signs of theoreticalphysics. How are we to get it using the path integral?
As explained in chapter I.3, the origin of (1) lies in the simple Gaussian integration
formula
/integraldisplay+∞
−∞dxe−1
2ax2=/radicalbigg
2π
a=√
2πe−1
2loga
Roughly speaking, we have to find a new type of integral so that the analog of the Gaussian
integral would go something like e+1
2loga.
126 | II. Dirac and the Spinor
Grassmann math
It turns out that the mathematics we need was invented long ago by Grassmann. Let
us postulate a new kind of number, called the Grassmann or anticommuting number,such that if ηandξare Grassmann numbers, then ηξ=−ξη. In particular, η
2=0.
Heuristically, this mirrors the anticommutation relation satisfied by the spinor field.Grassmann assumed that any function of ηcan be expanded in a Taylor series. Since η
2=0,
the most general function of ηisf( η)=a+bη, with aandbtwo ordinary numbers.
How do we define integration over η? Grassmann noted that an essential property of
ordinary integrals is that we can shift the dummy integration variable:/integraltext+∞
−∞dxf(x +c)=/integraltext+∞
−∞dxf(x) . Thus, we should also insist that the Grassmann integral obey the rule/integraltext
dηf(η +ξ)=/integraltext
dηf(η) , where ξis an arbitrary Grassmann number. Plugging into the
most general function given above, we find that/integraltext
dηbξ=0. Since ξis arbitrary this can only
hold if we define/integraltext
dηb=0 for any ordinary number b, and in particular/integraltext
dη≡/integraltext
dη1=0
Since given three Grassmann numbers χ,η, andξ, we have χ(ηξ) =(ηξ)χ , that is, the
product (ηξ) commutes with any Grassmann number χ, we feel that the product of two
anticommuting numbers should be an ordinary number. Thus, the integral/integraltext
dηη is just
an ordinary number that we can simply take to be 1: This fixes the normalization of dη.
Thus Grassmann integration is extraordinarily simple, being defined by two rules:
/integraldisplay
dη=0 (6)
and
/integraldisplay
dηη=1 (7)
With these two rules we can integrate any function of η:
/integraldisplay
dηf(η) =/integraldisplay
dη(a+bη)=b (8)
ifbis an ordinary number so that f( η) is Grassmannian, and
/integraldisplay
dηf(η) =/integraldisplay
dη(a+bη)=−b (9)
ifbis Grassmannian so that f( η) is an ordinary number. Note that the concept of a
range of integration does not exist for Grassmann integration. It is much easier to masterGrassmann integration than ordinary integration!
Letηand¯ηbe two independent Grassmann numbers and aan ordinary number. Then
the Grassmannian analog of the Gaussian integral gives
/integraldisplay
dη/integraldisplay
d¯ηe¯ηaη=/integraldisplay
dη/integraldisplay
d¯η(1+¯ηaη)=/integraldisplay
dηaη=a=e+loga(10)
Precisely what we had wanted!
We can generalize immediately: Let η=(η1,η2,...,ηN)beNGrassmann numbers,
and similarly for ¯η; we then have
II.5. Feynman Diagrams for Fermions | 127
/integraldisplay
dη/integraldisplay
d¯ηe¯ηAη=detA (11)
forA={Aij}an antisymmetric NbyNmatrix. (Note that contrary to the bosonic case, the
inverse of Aneed not exist.) We can further generalize to a functional integral.
As we will see shortly, we now have all the mathematics we need.
Grassmann path integral
In analogy with the generating functional for the scalar field
Z=/integraldisplay
DϕeiS(ϕ)=/integraldisplay
Dϕei/integraltext
d4x1
2[(∂ϕ)2−(m2−iε)ϕ2]
we would naturally write the generating functional for the spinor field as
Z=/integraldisplay
DψD ¯ψeiS(ψ ,¯ψ)=/integraldisplay
Dψ/integraldisplay
D¯ψei/integraltext
d4x¯ψ(i/negationslash∂−m+iε)ψ
Treating the integration variables ψand¯ψas Grassmann-valued Dirac spinors, we imme-
diately obtain
Z=/integraldisplay
Dψ/integraldisplay
D¯ψei/integraltext
d4x¯ψ(i/negationslash∂−m+iε)ψ=C/primedet(i/negationslash∂−m+iε)
=C/primeetr log(i/negationslash∂−m+iε)(12)
where C/primeis some multiplicative constant. Using the cyclic property of the trace, we note
that (here mis understood to be m−iε)
tr log(i/negationslash∂−m)=tr logγ5(i/negationslash∂−m)γ5=tr log(−i/negationslash∂−m)
=1
2[tr log(i/negationslash∂−m)+tr log(−i/negationslash∂−m)]
=1
2tr log(∂2+m2). (13)
Thus, Z=C/primee1
2tr log(∂2+m2−iε)[compare with (1)!].
We see that we get the same vacuum energy we obtained in chapter II.2 using the
canonical formalism if we remember that the trace operation here contains a factor of4 compared to the trace operation in (1), since (i/negationslash∂−m)i sa4b y4matrix.
Heuristically, we can now see the necessity for Grassmann variables. If we were to treat
ψand¯ψas complex numbers in (12), we would obtain something like (1/det[i/negationslash∂−m])=
e
−tr log (i/negationslash∂−m)and so have the wrong sign for the vacuum energy. We want the determinant
to come out in the numerator rather than in the denominator.
Dirac propagator
Now that we have learned that the Dirac field is to be quantized by a Grassmann path
integral we can introduce Grassmannian spinor sources ηand¯η:
Z(η ,¯η)=/integraldisplay
DψD ¯ψei/integraltext
d4x[¯ψ(i/negationslash∂−m)ψ +¯ηψ+¯ψη](14)
128 | II. Dirac and the Spinor
pp + kpk
Figure II.5.1
and proceed pretty much as before. Completing the square just as in the case of the scalar
field, we have
¯ψKψ +¯ηψ+¯ψη=(¯ψ+¯ηK−1)K(ψ +K−1η)−¯ηK−1η (15)
and thus
Z(η ,¯η)=C/prime/primee−i¯η(i/negationslash∂−m)−1η(16)
The propagator S(x) for the Dirac field is the inverse of the operator (i/negationslash∂−m): in other
words, S(x) is determined by
(i/negationslash∂−m)S(x) =δ(4)(x) (17)
As you can verify, the solution is
iS(x)=/integraldisplayd4p
(2π)4ie−ipx
/negationslashp−m+iε(18)
in agreement with (II.2.22).
Feynman rules for fermions
We can now derive the Feynman rules for fermions in the same way that we derived
the Feynman rules for a scalar field. For example, consider the theory of a scalar fieldinteracting with a Dirac field
L=¯ψ(iγμ∂μ−m)ψ+1
2[(∂ϕ)2−μ2ϕ2]−λϕ4+fϕ¯ψψ (19)
The generating functional
Z(η ,¯η,J)=/integraldisplay
DψD ¯ψDϕeiS(ψ ,¯ψ,ϕ)+i/integraltext
d4x(Jϕ+¯ηψ+¯ψη)(20)
can be evaluated as a double series in the couplings λandf. The Feynman rules (not
repeating the rules involving only the boson) are as follows:
1. Draw a diagram with straight lines for the fermion and dotted lines for the boson, and label
each line with a momentum, for example, as in figure II.5.1.
II.5. Feynman Diagrams for Fermions | 129
2. Associate with each fermion line the propagator
i
/negationslashp−m+iε=i/negationslashp+m
p2−m2+iε(21)
3. Associate with each interaction vertex the coupling factor ifand the factor (2π)4δ(4)(/summationtext
inp
−/summationtext
outp)expressing momentum conservation (the two sums are taken over the incoming
and the outgoing momenta, respectively).
4. Momenta associated with internal lines are to be integrated over with the measure/integraltext
[d4p/(2π)4].
5. External lines are to be amputated. For an incoming fermion line write u(p ,s)and for
an outgoing fermion line ¯u(p/prime,s/prime). The sources and sinks have to recognize the spin
polarization of the fermion being produced and absorbed. [For antifermions, we would have
¯v(p ,s)andv(p/prime,s/prime). You can see from (II.2.10) that an outgoing antifermion is associated
withvrather than ¯v.]
6. A factor of (−1)is to be associated with each closed fermion line. The spinor index carried
by the fermion should be summed over, thus leading to a trace for each closed fermion line.[For an example, see (II.7.7–9).]
Note that rule 6 is unique to fermions, and is needed to account for their negative
contribution to the vacuum energy. The Feynman diagram corresponding to vacuumfluctuation has no external line. I will discuss these points in detail in chapter IV .3.
For the theory of a massive vector field interacting with a Dirac field mentioned in
chapter II.1
L=¯ψ[iγμ(∂μ−ieAμ)−m]ψ−1
4FμνFμν+1
2μ2AμAμ(22)
the rules differ from above as follows. The vector boson propagator is given by
i
k2−μ2/parenleftbiggkμkν
μ2−gμν/parenrightbigg
(23)
and thus each vector boson line is associated not only with a momentum, but also with
indices μandν. The vertex (figure II.5.2) is associated with ieγμ.
p + k pk
ieγμ
Figure II.5.2
130 | II. Dirac and the Spinor
If the vector boson line in figure II.5.2 is external and on shell, we have to specify its
polarization. As discussed in chapter I.5, a massive vector boson has three degrees ofpolarizations described by the polarization vector ε
(a)
μfora=1, 2, 3. The amplitude for
emitting or absorbing a vector boson with polarization aisieγμε(a)
μ=ie/negationslashε(a)
μ.
In Schwinger’s sorcery, the source for producing a vector boson Jμ(x), in contrast to
the source for producing a scalar meson J(x) , carries a Lorentz index. Work in momen-
tum space. Current conservation kμJμ(k)=0 implies that we can decompose Jμ(k)=
/Sigma13
a=1J(a)(k)ε(a)
μ(k). The clever experimentalist sets up her machine, that is, chooses the
functions J(a)(k), so as to produce a vector boson of the desired momentum kand polar-
ization a. Current conservation requires kμε(a)
μ(k)=0. For kμ=(ω(k) ,0 ,0 ,k) , we could
choose
ε(1)
μ(k)=(0, 1, 0, 0 ),ε(2)
μ(k)=(0, 0, 1, 0 ),ε(3)
μ(k)=(−k ,0 ,0 ,ω(k))/m (24)
In the canonical formalism, we have in analogy with the expansion of the scalar field ϕ
in (I.8.11)
Aμ(/vectorx,t)=/integraldisplaydDk/radicalbig
(2π)D2ωk/Sigma13
a=1{a(a)(/vectork)ε(a)
μ(k)e−i(ω kt−/vectork./vectorx)+a(a)†(/vectork)ε(a)∗
μ(k)ei(ωkt−/vectork./vectorx)} (25)
(I trust you not to confuse the letter aused to denote annihilation and used to label
polarization.) The point is that in contrast to ϕ,Aμcarries a Lorentz index, which the
creation and annihilation operators have to “know about” (through the polarization label.)It is instructive to compare with the expansion of the fermion field ψin (II.2.10): the spinor
index αonψis carried in the expansion by the spinors u(p ,s)andv(p ,s). In each case,
an index (μ in the case of the vector and αin the case of the spinor) known to the Lorentz
group is “traded” for a label specifying the spin polarization (a andsrespectively.)
A minor technicality: notice that I have complex conjugated the polarization vector
associated with the creation operator a
(a)†(/vectork)in (25) even though the polarization vectors in
(24) are real. This is because experimentalists sometimes enjoy using circularly polarizedphotons with polarization vectors ε
(1)
μ(k)=(0, 1,i,0)/√
2,ε(2)
μ(k)=(0, 1,−i,0)/√
2.
pp + kpk
Figure II.5.3
II.5. Feynman Diagrams for Fermions | 131
Exercises
II.5.1 Write down the Feynman amplitude for the diagram in figure II.5.1 for the scalar theory (19). The answer
is given in chapter III.3.
II.5.2 Applying the Feynman rules for the vector theory (22) show that the amplitude for the diagram in
figure II.5.3 is given by
(ie)2i2/integraldisplayd4k
(2π)41
k2−μ2/parenleftbiggkμkν
μ2−gμν/parenrightbigg
¯u(p)γν/negationslashp+ /negationslashk+m
(p+k)2−m2γμu(p) (26)
II.6 Electron Scattering and Gauge Invariance
Electron-proton scattering
We will now finally calculate a physical process that experimentalists can go out and
measure. Consider scattering an electron off a proton. (For the moment let us ignore thestrong interaction that the proton also participates in. We will learn in chapter III.6 howto take this fact into account. Here we pretend that the proton, just like the electron, is astructureless spin-
1
2fermion obeying the Dirac equation.) To order e2the relevant Feynman
diagram is given in figure II.6.1 in which the electron and the proton exchange a photon.
But wait, from chapter I.5 we only know how to write down the propagator iDμν=
i/parenleftbigkμkν
μ2−ημν/parenrightbig
/(k2−μ2)for a hypothetical massive photon. (Trivial notational change: the
mass of the photon is now called μ, since mis reserved for the mass of the electron and
Mfor the mass of the proton.) In that chapter I outlined our philosophy: we will plunge
ahead and calculate with a nonzero μand hope that at the end we can set μto zero. Indeed,
when we calculated the potential energy between two external charges, we find that we canletμ→0 without any signs of trouble [see (I.5.6)]. In this chapter and the next, we would
like to see whether this will always be the case.
Applying the Feynman rules, we obtain the amplitude for the diagram in figure II.6.1
(withk=P−pthe momentum transfer in the scattering)
M(P,PN)=(−ie)(ie)i
(P−p)2−μ2/parenleftbiggkμkν
μ2−ημν/parenrightbigg
¯u(P)γμu(p)¯u(PN)γνu(pN) (1)
We have suppressed the spin labels and used the subscript N(for nucleon) to refer to the
proton.
Now notice that
kμ¯u(P)γμu(p)=(P−p)μ¯u(P)γμu(p)=¯u(P)(/negationslashP− /negationslashp)u(p) =¯u(P)(m −m)u(p) =0 (2)
by virtue of the equations of motion satisfied by ¯u(P) andu(p) . Similarly, kμ¯u(PN)γμu(pN)
=0.
II.6. Scattering and Gauge Invariance | 133
kPN
pNP
p
Figure II.6.1
This important observation implies that the kμkν/μ2term in the photon propagator does
not enter. Thus
M(P,PN)=−ie2 1
(P−p)2−μ2¯u(P)γμu(p)¯u(PN)γμu(pN) (3)
and we can now set the photon mass μto zero with impunity and replace (P−p)2−μ2
in the denominator by (P−p)2.
Note that the identity that allows us to set μto zero is just the momentum space version
of electromagnetic current conservation ∂μJμ=∂μ(¯ψγμψ)=0. You would notice that
this calculation is intimately related to the one we did in going from (I.5.4) to (I.5.5), with
¯u(P)γμu(p) playing the role of Jμ(k).
Potential scattering
That the proton mass Mis so much larger than the electron mass mallows us to make
a useful approximation familiar from elementary physics. In the limit M/m tending to
infinity, the proton hardly moves, and we could use, for the proton, the spinors for a particleat rest given in chapter II.2, so that ¯u(P
N)γ0u(pN)≈1 and ¯u(PN)γiu(pN)≈0. Thus
M=−ie2
k2¯u(P)γ0u(p) (4)
We recognize that we are scattering the electron in the Coulomb potential generated by
the proton. Work out the (familiar) kinematics: p=(E,0 ,0 ,|/vector p|)andP=(E,0 ,|/vectorp|sinθ,
|/vectorp|cosθ). We see that k=P−pis purely spacelike and k2=−/vectork2=− 4|/vectorp|2sin2(θ/2).
Recall from (I.4.7) that
/integraldisplay
d3xei/vectork./vectorx/parenleftbigg
−e
4πr/parenrightbigg
=−e
/vectork2(5)
We represent potential scattering by the Feynman diagram in figure II.6.2: the proton has
disappeared and been replaced by a cross, which supplies the virtual photon the electroninteracts with. It is in this sense that you could think of the Coulomb potential picturesquelyas a swarm of virtual photons.
134 | II. Dirac and the Spinor
k
XP
p
Figure II.6.2
Once again, it is instructive to use the canonical formalism to derive this expression for
Mfor potential scattering. We want the transition amplitude /angbracketleftP,S|e−iHT|p,s/angbracketrightwith the
single electron state |p,s/angbracketright≡b†(p,s)|0/angbracketright. The term in the Lagrangian describing the elec-
tron interacting with the external c-number potential Aμ(x) is given in (II.5.22) and thus
to leading order we have the transition amplitude ie/integraltext
d4x/angbracketleftP,S|¯ψ(x)γμψ(x)|p,s/angbracketrightAμ(x).
Using (II.2.10–11) we evaluate this as
ie/integraldisplay
d4x(1/ρ(P))(1 /ρ(p))( ¯u(P ,S)γμu(p ,s))ei(P−p)xAμ(x)
Hereρ(p) denotes the fermion normalization factor/radicalBig
(2π)3Ep/m in (II.2.10). Given that
the Coulomb potential has only a time component and does not depend on time, we see
that integration over time gives us an energy conservation delta function, and integrationover space the Fourier transform of the potential, as in (5). Thus the above becomes
(1/ρ(P))(1 /ρ(p))(2 π)δ(E
P−Ep)(−ie2
/vectork2)¯u(P ,S)γ0u(p ,s). Satisfyingly, we have recovered
the Feynman amplitude up to normalization factors and an energy conservation deltafunction, just as in (I.8.16) except for the substitution of boson for fermion normalizationfactors. Notice that we have energy conservation but not 3-momentum conservation, a factwe understand perfectly well when we dribble a basketball ball, for example.
Electron-electron scattering
Next, we graduate to two electrons scattering off each other: e−(p1)+e−(p2)→e−(P1)+
e−(P2). Here we have a new piece of physics: the two electrons are identical. A profound
tenet of quantum physics states that we cannot distinguish between the two outgoing
electrons. Now there are two Feynman diagrams (see fig. II.6.3) to order e2, obtained
by interchanging the two outgoing electrons. The electron carrying momentum P1could
have “come from” the incoming electron carrying momentum p1or the incoming electron
carrying momentum p2.
We have for figure II.6.3a the amplitude
A(P 1,P2)=(ie2/(P 1−p1)2)¯u(P 1)γμu(p 1)¯u(P 2)γμu(p 2)
II.6. Scattering and Gauge Invariance | 135
(a)P1 P2
p1 p2k
(b)P1 P2
p1 p2
Figure II.6.3
as before. We have only indicated the dependence of Aon the final momenta, suppressing
the other dependence. By Fermi statistics, the amplitude for the diagram in figure II.6.3is then −A(P
2,P1). Thus the invariant amplitude for two electrons of momentum p1and
p2to scatter into two electrons with momentum P1andP2is
M=A(P 1,P2)−A(P 2,P1) (6)
To obtain the cross section we have to square the amplitude
|M|2=[|A(P 1,P2)|2+(P1↔P2)]−2R eA(P 2,P1)∗A(P 1,P2) (7)
136 | II. Dirac and the Spinor
At this point we have to do a fair amount of arithmetic, but keep in mind that there
is nothing conceptually intricate in what follows. First, we have to learn to complex con-jugate spinor amplitudes. Using (II.1.15), note that in general (¯u(p
/prime)γμ...γνu(p))∗=
u(p)†γ†
ν...γ†
μγ0u(p/prime)=¯u(p)γν...γμu(p/prime). Here γμ...γνrepresents a product of any
number of γmatrices. Complex conjugation reverses the order of the product and inter-
changes the two spinors. Thus we have
|A(P 1,P2)|2=e4
k4[¯u(P 1)γμu(p 1)¯u(p 1)γνu(P1)][¯u(P 2)γμu(p2)¯u(p 2)γνu(P 2)] (8)
which factorizes with one factor involving spinors carrying momentum with subscript 1
and another factor involving spinors carrying momentum with subscript 2. In contrast,the interference term A(P
2,P1)∗A(P1,P2)does not factorize.
In the simplest experiments, the initial electrons are unpolarized, and the polarization
of the outgoing electrons is not measured. We average over initial spins and sum over finalspins using (II.2.8):
/summationdisplay
su(p ,s)¯u(p ,s)=/negationslashp+m
2m(9)
In averaging and summing |A(P1,P2)|2we encounter the object (displaying the spin labels
explicitly)
τμν(P1,p1)≡1
2/summationdisplay/summationdisplay
¯u(P 1,S)γμu(p 1,s)¯u(p 1,s)γνu(P 1,S) (10)
=1
2(2m)2tr(/negationslashP1+m)γμ(/negationslashp1+m)γν(11)
which is to be multiplied by τμν(P2,p2).
Well, we, or rather you, have to develop some technology for evaluating the trace of
products of gamma matrices. The key observation is that the square of a gamma matrixis either +1o r −1, and different gamma matrices anticommute. Clearly, the trace of a
product of an odd number of gamma matrices vanish. Furthermore, since there are onlyfour different gamma matrices, the trace of a product of six gamma matrices can always bereduced to the trace of a product of four gamma matrices, since there are always pairs ofgamma matrices that are equal and can be brought together by anticommuting. Similarlyfor the trace of a product of an even higher number of gamma matrices.
Hence τ
μν(P1,p1)=1
2(2m)2(tr(/negationslashP1γμ/negationslashp1γν)+m2tr(γμγν)). Writing tr (/negationslashP1γμ/negationslashp1γν)=
P1ρp1λtr(γργμγλγν)and using the expressions for the trace of a product of an even
number of gamma matrices listed in appendix D, we obtain τμν(P1,p1)=1
2(2m)24(Pμ
1pν
1−
ημνP1.p1+Pν
1pμ
1+m2ημν).
In averaging and summing A(P2,P1)∗A(P 1,P2)we encounter the more involved object
κ≡1
22/summationdisplay/summationdisplay/summationdisplay/summationdisplay
¯u(P 1)γμu(p 1)¯u(P 2)γμu(p 2)¯u(p 1)γνu(P 2)¯u(p 2)γνu(P 1) (12)
where for simplicity of notation we have suppressed the spin labels. Applying (9) we can
writeκas a single trace. The evaluation of κis quite tedious, since it involves traces of
products of up to eight gamma matrices.
II.6. Scattering and Gauge Invariance | 137
We will be content to study electron-electron scattering in the relativistic limit in which
mmay be neglected compared to the momenta. As explained in chapter II.2, while we
are using the “rest normalization” for spinors we can nevertheless set mto 0 wherever
possible. Then
κ=1
4(2m)4tr(/negationslashP1γμ/negationslashp1γν/negationslashP2γμ/negationslashp2γν) (13)
Applying the identities in appendix D to (13) we obtain tr (/negationslashP1γμ/negationslashp1γν/negationslashP2γμ/negationslashp2γν)=
−2tr(/negationslashP1γμ/negationslashp1/negationslashp2γμ/negationslashP2)=− 32p1.p2P1.P2.
In the same limit τμν(P1,p1)=2
(2m)2(Pμ
1pν
1+Pν
1pμ
1−ημνP1.p1)and thus
τμν(P1,p1)τμν(P2,p2)=4
(2m)4(Pμ
1pν
1+Pν
1pμ
1−ημνP1.p1)(2P2μp2ν−ημνP2.p2) (14)
=4.2
(2m)4(p1.p2P1.P2+p1.P2p2.P1) (15)
An amusing story to break up this tedious calculation: Murph Goldberger, who was
a graduate student at the University of Chicago after working on the Manhattan Projectduring the war and whom I mentioned in chapter II.1 regarding the Feynman slash, toldme that Enrico Fermi marvelled at this method of taking a trace that young people wereusing to sum over spin-
1
2polarizations. Fermi and others in the older generation had
simply memorized the specific form of the spinors in the Dirac basis (which you know fromdoing exercise II.1.3) and consequently the expressions for ¯u(P ,S)γ
μu(p ,s). They simply
multiplied these expressions together and added up the different possibilities. Fermi wasskeptical of the fancy schmancy method the young Turks were using and challenged Murphto a race on the blackboard. Of course, with his lightning speed, Fermi won. To me, it isamazing, living in the age of string theory, that another generation once regarded thetrace as fancy math. I confessed that I was even a bit doubtful of this story until I looked atFeynman’s book Quantum Electrodynamics, but guess what, Feynman indeed constructed,
on page 100 in the edition I own, a table showing the result for the amplitude squared forvarious spin polarizations. Some pages later, he mentioned that polarizations could alsobe summed using the spur (the original German word for trace). Another amusing aside:spur is cognate with the English word spoor, meaning animal droppings, and hence alsomeaning track, trail, and trace. All right, back to work!
While it is not the purpose of this book to teach you to calculate cross sections for a
living, it is character building to occasionally push calculations to the bitter end. Hereis a good place to introduce some useful relativistic kinematics. In calculating the crosssection for the scattering process p
1+p2→P1+P2(with the masses of the four particles
all different in general) we typically encounter Lorentz invariants such as p1.P2. A priori,
you might think there are six such invariants, but in fact, you know that there are only
physical variables, the incident energy Eand the scattering angle θ. The cleanest way to
organize these invariants is to introduce what are called Mandelstam variables:
s≡(p1+p2)2=(P1+P2)2(16)
t≡(P1−p1)2=(P2−p2)2(17)
u≡(P2−p1)2=(P1−p2)2(18)
138 | II. Dirac and the Spinor
You know that there must be an identity reducing the three variables s,t, anduto two.
Show that (with an obvious notation)
s+t+u=m2
1+m2
2+M2
1+M2
2(19)
For our calculation here, we specialize to the center of mass frame in the relativistic limit
p1=E(1, 0, 0, 1 ),p2=E(1, 0, 0, −1),P1=E(1, sinθ,0 ,c o s θ), andP2=E(1,−sinθ,0 ,
−cosθ). Hence
p1.p2=P1.P2=2E2=1
2s (20)
p1.P1=p2.P2=2E2sin2θ
2=−1
2t (21)
and
p1.P2=p2.P1=2E2cos2θ
2=−1
2u (22)
Also, in this limit (P1−p1)4=(−2p1.P1)2=16E4sin4(θ/2)=t4. Putting it together, we
obtain1
4/summationtext/summationtext/summationtext/summationtext|M|2=(e4/4m4)f (θ), where
f( θ)=s2+u2
t2+2s2
tu+s2+t2
u2
=s4+t4+u4
t2u2
=/parenleftBigg
1+cos4(θ/2)
sin4(θ/2)+2
sin2(θ/2)cos2(θ/2)+1+sin4(θ/2)
cos4(θ/2)/parenrightBigg
(23)
=2/parenleftbigg1
sin4(θ/2)+1+1
cos4(θ/2)/parenrightbigg
(24)
The physical origin of each of the terms in (23) [before we simplify with trigonometric
identities] to get to (24) is clear. The first term strongly favors forward scattering dueto the photon propagator ∼1/k
2blowing up at k∼0. The third term is required by the
indistinguishability of the two outgoing electrons: the scattering must be symmetric underθ→π−θ, since experimentalists can’t tell whether a particular incoming electron has
scattered forward or backward. The second term is the most interesting of all: it comes fromquantum interference. If we had mistakenly thought that electrons are bosons and takenthe plus sign in (6), the second term in f( θ) would come with a minus sign. This makes
a big difference: for instance, f( π/ 2)would be 5 −8+5=2 instead of 5 +8+5=18.
Since the conversion of a squared probability amplitude to a cross section is conceptually
the same as in nonrelativistic quantum mechanics (divide by the incoming flux, etc.), Iwill relegate the derivation to an appendix and let you go the last few steps and obtain thedifferential cross section as an exercise:
dσ
d/Omega1=α2
8E2f( θ) (25)
with the fine structure constant α≡e2/4π≈1/137.
II.6. Scattering and Gauge Invariance | 139
An amazing subject
When you think about it, theoretical physics is truly an amazing business. After the
appropriate equipments are assembled and high energy electrons are scattered off eachother, experimentalists indeed would find the differential cross section given in (25). Thereis almost something magical about it.
Appendix: Decay rate and cross section
To make contact with experiments, we have to convert transition amplitudes into the scattering cross sections
and decay rates that experimentalists actually measure. I assume that you are already familiar with the physical
concepts behind these measurements from a course on nonrelativistic quantum mechanics, and thus here wefocus more on those aspects specific to quantum field theory.
To be able to count states, we adopt an expedient probably already familiar to you from quantum statistical
mechanics, namely that we enclose our system in a box, say a cube with length Lon each side with Lmuch larger
than the characteristic size of our system. With periodic boundary conditions, the allowed plane wave states e
i/vectorp./vectorx
carry momentum
/vectorp=2π
L(nx,ny,nz) (26)
where the ni’s are three integers. The allowed values of momentum form a lattice of points in momentum space
with spacing 2 π/L between points. Experimentalists measure momentum with finite resolution, small but much
larger than 2 π/L . Thus, an infinitesimal volume d3pin momentum space contains d3p/(2π/L)3=Vd3p/(2π)3
states with V=L3the volume of the box. We obtain the correspondence
/integraldisplayd3p
(2π)3f( p)↔1
V/summationdisplay
pf( p) (27)
In the sum the values of pranges over the discrete values in (26). The correspondence (27) between continuum
normalization and the discrete box normalization implies that
δ(3)(/vectorp−/vectorp/prime)↔V
(2π)3δ/vectorp/vectorp/prime (28)
with the Kronecker delta δ/vectorp/vectorp/primeequal to 1 if /vectorp=/vectorp/primeand 0 otherwise. One way of remembering these correspondences
is simply by dimensional matching.
Let us now look at the expansion (I.8.17) of a complex scalar field
ϕ(/vectorx,t)=/integraldisplayd3k/radicalbig
(2π)32ωk[a(/vectork)e−i(ω kt−/vectork./vectorx)+b†(/vectork)ei(ωkt−/vectork./vectorx)] (29)
in terms of creation and annihilation operators. Henceforth, in order not to clutter up the page, I will abuse
notation slightly, for example, dropping the arrows on vectors when there is no risk of confusion. Going over to
the box normalization, we replace the commutation relation [ a(k) ,a†(k/prime)]=δ(3)(/vectork−/vectork/prime)by
[a(k) ,a†(k/prime)]=V
(2π)3δ/vectork/vectork/prime (30)
We now normalize the creation and annihilation operators by
a(k)=/parenleftbiggV
(2π)3/parenrightbigg1
2
˜a(k) (31)
140 | II. Dirac and the Spinor
so that
[˜a(k) ,˜a†(k/prime)]=δk,k/prime (32)
Thus the state |/vectork/angbracketright≡˜a†(/vectork)|0/angbracketrightis properly normalized: /angbracketleft/vectork|/vectork/angbracketright=1. Using (27) and (31), we end up with
ϕ(x)=1
V1
2/summationdisplay
k1/radicalbig
2ωk(˜ae−ikx+˜b†eikx) (33)
We specified a complex, rather than a real, scalar field because then, as you showed in exercise I.8.4, a
conserved current can be defined with the corresponding charge Q=/integraltext
d3xJ0=/integraltext
d3k(a†(k)a(k) −b†(k)b(k)) →/summationtext
k(˜a†(k)˜a(k)−˜b†(k)˜b(k)) . It follows immediately that /angbracketleft/vectork|Q|/vectork/angbracketright=1, so that for the state |/vectork/angbracketrightwe have one particle
in the box.
To derive the formula for the decay rate, we focus, for the sake of pedagogical clarity, on a toy Lagrangian
L=g(η†ξ†ϕ+h.c.)describing the decay ϕ→η+ξof a meson into two other mesons. (As usual, we display
only the part of the Lagrangian that is of immediate interest. In other words, we suppress the stuff you have longsince mastered: L=∂ϕ
†∂ϕ−mϕϕ†ϕ+...and all the rest.)
The transition amplitude /angbracketleft/vectorp,/vectorq|e−iHT|/vectork/angbracketrightis given to lowest order by A=i/angbracketleft/vectorp,/vectorq|/integraltext
d4x(gη†(x)ξ†(x)ϕ(x)) |/vectork/angbracketright.
Here we use the states we “carefully” normalized above, namely the ones created by the various “analogs” of ˜a†.
(Just as in quantum mechanics, strictly, we should use wave packets instead of plane wave states. I assume thatyou have gone through that at least once.) Plugging in the various “analogs” of (33), we have
A=ig(1
V1
2)3/summationdisplay
p/prime/summationdisplay
q/prime/summationdisplay
k/prime1/radicalbig2ωp/prime2ωq/prime2ωk/prime/integraldisplay
d4xei(p/prime+q/prime−k/prime)/angbracketleft/vectorp,/vectorq|˜a†(p/prime)˜a†(q/prime)˜a(k/prime)|/vectork/angbracketright
=ig1
V3
21/radicalbig2ωp2ωq2ωk(2π)4δ(4)(p+q−k)(34)
Here we have committed various minor transgressions against notational consistency. For example, since
the three particles ϕ,η, andξhave different masses, the symbol ωrepresents, depending on context, different
functions of its subscript (thus ωp=/radicalBig
/vectorp2+m2
η, and so forth). Similarly, ˜a(k/prime)should really be written as ˜aϕ(k/prime),
and so forth. Also, we confound 3- and 4-momenta. I would like to think that these all fall under the category of
what the Catholic church used to call venial sins. In any case, you know full well what I am talking about.
Next, we square the transition amplitude Ato find the transition probability. You might be worried, because
it appears that we will have to square the Dirac delta function. But fear not, we have enclosed ourselves in a box.
Furthermore, we are in reality calculating /angbracketleft/vectorp,/vectorq|e−iHT|/vectork/angbracketright, the amplitude for the state |/vectork/angbracketrightto become the state
|/vectorp,/vectorq/angbracketrightafter a large but finite time T. Thus we could in all comfort write
[(2π)4δ(4)(p+q−k)]2=(2π)4δ(4)(p+q−k)/integraldisplay
d4xei(p+q−k)x
=(2π)4δ(4)(p+q−k)/integraldisplay
d4x=(2π)4δ(4)(p+q−k)VT(35)
Thus the transition probability per unit time, aka the transition rate, is equal to
|A|2
T=V
V3/parenleftBigg
1
2ωp2ωq2ωk/parenrightBigg
(2π)4δ(4)(p+q−k)g2(36)
Recall that there are Vd3p/(2π)3states in the volume d3pin momentum space. Hence, multiplying the number
of final states (V d3p/(2π)3)(V d3q/(2π)3)by the transition rate |A|2/T, we obtain the differential decay rate of
a meson into two mesons carrying off momenta in some specified range d3pandd3q:
d/Gamma1=1
2ωkV
V3/parenleftBigg
Vd3p
(2π)32ωp/parenrightBigg/parenleftBigg
Vd3q
(2π)32ωq/parenrightBigg
(2π)4δ(4)(p+q−k)g2(37)
Yes sir, indeed, the factors of Vcancel, as they should.
To obtain the total decay rate /Gamma1we integrate over d3pandd3q. Notice the factor 1 /2ωk: the decay rate for a
moving particle is smaller than that of a resting particle by a factor m/ωk. We have derived time dilation, as we
had better.
II.6. Scattering and Gauge Invariance | 141
We are now ready to generalize to the decay of a particle carrying momentum Pintonparticles carrying
momenta k1,...,kn. For definiteness, we suppose that these are all Bose particles. First, we draw all the relevant
Feynman diagrams and compute the invariant amplitude M. (In our toy example, M=ig.) Second, the transition
probability contains a factor 1 /Vn+1, one factor of 1 /V for each particle, but when we squared the momentum
conservation delta function we also obtained a factor of VT, which converts the transition probability into a
transition rate and knocks off one power of V, leaving the factor 1 /Vn. Next, when we sum over final states, we
have a factor Vd3ki/((2π)32ωki)for each particle in the final state. Thus the factors of Vindeed cancel.
The differential decay rate of a boson of mass Min its rest frame is thus given by
d/Gamma1=1
2Md3k1
(2π)32ω(k1)...d3kn
(2π)32ω(kn)(2π)4δ(4)/parenleftBigg
P−n/summationdisplay
i=1ki/parenrightBigg
|M|2(38)
At this point, we recall that, as explained in chapter II.2, in the expansion of a fermion field into creation and
annihilation operators [see (II.2.10)], we have a choice of two commonly used normalizations, trivially related
by a factor (2m)1
2. If you choose to use the “rest normalization" so that spinors come out nice in the rest frame,
then the field expansion contains the normalization factor (Ep/m)1
2instead of the factor (2ωk)1
2for a Bose
field [see (I.8.11)]. This entails the trivial replacement, for each fermion, of the factor 2 ω(k)=2/radicalbig
/vectork2+m2by
E(p)/m =/radicalbig
/vectorp2+m2/m. In particular, for the decay rate of a fermion the factor 1 /2Mshould be removed. If
you choose the “any mass renormalization,” you have to remember to normalize the spinors appearing in M
correctly, but you need not touch the phase space factors derived here.
We next turn to scattering cross sections. As I already said, the basic concepts involved should already be
familiar to you from nonrelativistic quantum mechanics. Nevertheless, it may be helpful to review the basicnotions involved. For the sake of definiteness, consider some happy experimentalist sending a beam of haplesselectrons crashing into a stationary proton. The flux of the beam is defined as the number of electrons crossingan imagined unit area per unit time and is thus given by F=nv, where nandvdenote the density and velocity
of the electrons in the beam. The measured event rate divided by the flux of the beam is defined to be the crosssection σ, which has the dimension of an area and could be thought of as the effective size of the proton as seen
by the electrons.
It may be more helpful to go to the rest frame of the electrons, in which the proton is plowing through the
cloud of electrons like a bulldozer. In time /Delta1tthe proton moves through a distance v/Delta1t and thus sweeps through
a volume σv/Delta1t , which contains nσv/Delta1t electrons. Dividing this by /Delta1tgives us the event rate nvσ .
To measure the differential cross section, the experimentalist sets up, typically in the lab frame in which the
target particle is at rest, a detector spanning a solid angle d/Omega1=sinθdθdφ and counts the number of events per
unit time.
All of this is familiar stuff. Now we could essentially take over our calculation of the differential decay rate
almost in its entirety to calculate the differential cross section for the process p
1+p2→k1+k2+...+kn. With
two particles in the initial state we now have a factor of (1/V)n+2in the transition probability. But as before, the
square of the momentum conservation delta function produces one power of Vand counting the momentum
final states gives a factor Vn, so that we are left with a factor of 1 /V. You might be worried about this remaining
factor of 1 /V, but recall that we still have to divide by the flux, given by |/vectorv1−/vectorv2|n. Since we have normalized to
one particle in the box the density nis 1/V. Once again, all factors of Vcancel, as they must.
The procedure is thus to draw all relevant diagrams to the order desired and calculate the Feynman amplitude
Mfor the process p1+p2→k1+k2+...+kn. Then the differential cross section is given by (again assuming
all particles to be bosons)
dσ=1
|/vectorv1−/vectorv2|2ω(p 1)2ω(p 2)d3k1
(2π)32ω(k 1)...d3kn
(2π)32ω(kn)(2π)4δ(4)/parenleftBigg
p1+p2−n/summationdisplay
i=1ki/parenrightBigg
|M|2
(39)
We are implicitly working in a collinear frame in which the velocities of the incoming particles, /vectorv1and/vectorv2,
point in opposite directions. This class of frames includes the familiar center of mass frame and the lab frame (inwhich /vectorv
2=0). In a collinear frame, p1=E1(1, 0, 0, v1)andp2=E2(1, 0, 0, v2), and a simple calculation shows
that((p 1p2)2−m2
1m22)=(E1E2(v1−v2))2. We could write the factor |/vectorv1−/vectorv2|E1E2indσin the more invariant-
looking form ((p1p2)2−m2
1m22)1
2, thus showing explicitly that the differential cross section is invariant under
Lorentz boosts in the direction of the beam, as physically must be the case.
142 | II. Dirac and the Spinor
An often encountered case involves two particles scattering into two particles in the center of mass frame. Let
us do the phase space integral/integraltext
(d3k1/2ω1)(d3k2/2ω2)δ(4)(P−k1−k2)here for easy reference. We will do it in
two different ways for your edification.
We could immediately integrate over d3k2thus knocking out the 3-dimensional momentum conservation
delta function δ3(/vectork1+/vectork2). Writing d3k1=k2
1dk1d/Omega1, we integrate over the remaining energy conservation delta
function δ(/radicalBig
k2
1+m2
1+/radicalBig
k2
1+m2
2−Etotal). Using (I.2.13), we find that the integral over k1givesk1ω1ω2/E total,
where ω1≡/radicalBig
k2
1+m2
1andω2≡/radicalBig
k2
1+m2
2, withk1determined by/radicalBig
k2
1+m2
1+/radicalBig
k2
1+m2
2=Etotal. Thus we obtain
/integraldisplayd3k1
2ω1d3k2
2ω2δ(4)(P−k1−k2)=k1
4Etotal/integraldisplay
d/Omega1 (40)
Once again, if you use the “rest normalization" for fermions, remember to make the replacement as explained
above for the decay rate. The factor of1
4should be replaced by mf/2 for one fermion and one boson, and by
m1m2for two fermions.
Alternatively, we use (I.8.14) and regressing, write d3k2=/integraltext
d4k2θ(k0
2)δ(k2
2−m2
2)2ω2. Integrate over d4k2and
knock out the 4-dimensional delta function, leaving us with/integraltext
2ω2dk1k2
1d/Omega1/(2 ω12ω2)δ((P −k1)2−m2
2). The
argument of the delta function is E2
total−2Etotalk1+m2
1−m2
2, and thus integrating over k1we get a factor of
2Etotal in the denominator, giving a result in agreement with (40).
For the record, you could work out the kinematics and obtain
k1=/radicalBig
(E2
total−(m 1+m2)2)(E2
total−(m1−m2)2)/2Etotal
Evidently, this phase space integral also applies to the decay into two particles in the rest frame of the parent
particle, in which case we replace Etotal byM. In particular, for our toy example, we have
/Gamma1=g2
16πM3/radicalbig
(M2−(m+μ)2)(M2−(m−μ)2) (41)
The differential cross section for two-into-two scattering in the center of mass frame is given by
dσ
d/Omega1=1
(2π)2|/vectorv1−/vectorv2|2ω(p 1)2ω(p 2)k1
EtotalF|M|2(42)
In particular, in the text we calculated electron-electron scattering in the relativistic limit. As shown there, we
can write |M|2=|/hatwiderM|2/(2m)4in terms of some reduced invariant amplitude /hatwiderM. The factor 1 /(2m)4transforms
the factors 2 ω(p) into 2E. Things simplify enormously, with |/vectorv1−/vectorv2|=2 and k1=1
2Etotal, so that finally
dσ
d/Omega1=1
24(4π)2E2|/hatwiderM|2(43)
Last, we come to the statistical factor Sthat must be included in calculating the total decay rate and the total
cross section to avoid over-counting if there are identical particles in the final state. The factor Shas nothing to do
with quantum field theory per se and should already be familiar to you from nonrelativistic quantum mechanics.The rule is that if there are n
iidentical particles of type iin the final state, the total decay rate or the total cross
section must be multiplied by S=/Pi1i1/ni! to account for indistinguishability.
To see the necessity for this factor, it suffices to think about the simplest case of two identical Bose particles.
To be specific, consider electron-positron annihilation into two photons (which we will study in chapter II.8). Forsimplicity, average and sum over all spin polarizations. Let us calculate dσ/d/Omega1 according to (43) above. This is
the probability that a photon will check into a detector set up at angles θandφrelative to the beam direction. If
the detector clicks, then we know that the other photon emerged at an angle π−θrelative to the beam direction.
Thus the total cross section should be
σ=1
2/integraldisplay
d/Omega1dσ
d/Omega1=1
2/integraldisplayπ
0dθdσ
dθ(44)
(The second equality is for all the elementary cases we will encounter in which dσ does not depend on the
azimuthal angle φ.) In other words, to avoid double counting, we should divide by 2 if we integrate over the full
angular range of θ.
More formally, we argue as follows. In quantum mechanics, a set of states |α/angbracketrightis complete if 1 =/summationtext
α|α/angbracketleft/angbracketrightα|
(“decomposition of 1”). Acting with this on |β/angbracketrightwe see that these states must be normalized according to
/angbracketleftα|β/angbracketright=δαβ.
II.6. Scattering and Gauge Invariance | 143
Now consider the state
|k1,k2/angbracketright≡1√
2˜a†(k1)˜a†(k2)|0/angbracketright=|k2,k1/angbracketright (45)
containing two identical bosons. By repeatedly using the commutation relation (32), we compute /angbracketleftq1,q2|k1,k2/angbracketright=
/angbracketleft0|˜a(q 1)˜a(q 2)˜a†(k1)˜a†(k2)|0/angbracketright=1
2(δq1k1δq2k2+δq2k1δq1k2). Thus/summationtext
q1/summationtext
q2|q1,q2/angbracketright/angbracketleftq 1,q2|k1,k2/angbracketright
=/summationtext
q1/summationtext
q2|q1,q2/angbracketright1
2(δq1k1δq2k2+δq2k1δq1k2)=1
2(|k1,k2/angbracketright+|k2,k1/angbracketright)=|k1,k2/angbracketright. Thus the states |k1,k2/angbracketrightare nor-
malized properly. In the sum over states, we have 1 =...+/summationtext
q1/summationtext
q2|q1,q2/angbracketright/angbracketleftq 1,q2|+ ....
In other words, if we are to sum over q1andq2independently, then we must normalize our states as in (45)
with the factor of 1 /√
2. But then this factor would appear multiplying M. In calculating the total decay rate
or the total cross section, we are effectively summing over a complete set of final states. In summary, we havetwo options: either we treat the integration over d
3k1d3k2as independent in which case we have to multiply the
integral by1
2, or we integrate over only half of phase space.
We readily generalize from this factor of1
2to the statistical factor S.
In closing, let me mention two interesting pieces of physics.To calculate the cross section σ, we have to divide by the flux, and hence σis proportional to 1 /|/vectorv
1−/vectorv2|.
For exothermal processes, such as electron-positron annihilation into photons or slow neutron capture, σcould
become huge as the relative velocity vrel→0. Fermi exploited this fact to great advantage in studying nuclear
fission. Note that although the cross section, which has dimension of an area, formally goes to infinity, thereaction rate (the number of reactions per unit time) remains finite.
Positronium decay into photons is an example of a bound state decaying in finite time. In positronium, the
positron and electron are not approaching each other in plane wave states, as we assumed in our cross sectioncalculation. Rather, the probability (per unit volume) that the positron finds itself near the electron is given by|ψ(0)|
2according to elementary quantum mechanics, with ψ(x) the bound state wave function for whatever
state of positronium we are interested in. In other words, |ψ(0)|2gives the volume density of positrons near the
electron. Since vσis a volume divided by time, the decay rate is given by /Gamma1=vσ|ψ(0)|2.
Exercises
II.6.1 Show that the differential cross section for a relativistic electron scattering in a Coulomb potential is
given by
dσ
d/Omega1=α2
4/vectorp2v2sin4(θ/2)(1−v2sin2(θ/2)).
Known as the Mott cross section, it reduces to the Rutherford cross section you derived in a course on
quantum mechanics in the limit the electron velocity v→0.
II.6.2 To order e2the amplitude for positron scattering off a proton is just minus the amplitude (3) for electron
scattering off a proton. Thus, somewhat counterintuitively, the differential cross sections for positronscattering off a proton and for electron scattering off a proton are the same to this order. Show that to
the next order this is no longer true.
II.6.3 Show that the trace of a product of odd number of gamma matrices vanishes.
II.6.4 Prove the identity s+t+u=/summationtext
am2
a.
II.6.5 Verify the differential cross section for relativistic electron electron scattering given in (25).
II.6.6 For those who relish long calculations, determine the differential cross section for electron-electron
scattering without taking the relativistic limit.
II.6.7 Show that the decay rate for one boson of mass Minto two bosons of masses mandμis given by
/Gamma1=|M|2
16πM3/radicalbig
(M2−(m+μ)2)(M2−(m−μ)2)
II.7 Diagrammatic Proof of Gauge Invariance
Gauge invariance
Conceptually, rather than calculate cross sections, we have the more important task of
proving that we can indeed set the photon mass μequal to zero with impunity in calculating
any physical process. With μ=0, the Lagrangian given in chapter II.1 becomes the
Lagrangian for quantum electrodynamics:
L=¯ψ[iγμ(∂μ−ieAμ)−m]ψ−1
4FμνFμν(1)
We are now ready for one of the most important observations in the history of theoretical
physics. Behold, the Lagrangian is left invariant by the gauge transformation
ψ(x)→ei/Lambda1(x)ψ(x) (2)
and
Aμ(x)→Aμ(x)+1
iee−i/Lambda1(x)∂μei/Lambda1(x)=Aμ(x)+1
e∂μ/Lambda1(x) (3)
which implies
Fμν(x)→Fμν(x) (4)
You are of course already familiar with (3) and the invariance of Fμνfrom classical
electromagnetism.
In contemporary theoretical physics, gauge invariance1is regarded as fundamental and
all important, as we will see later. The modern philosophy is to look at (1) as a consequence
of (2) and (3). If we want to construct a gauge invariant relativistic field theory involving aspin
1
2and a spin 1 field, then we are forced to quantum electrodynamics.
1The discovery of gauge invariance was one of the most arduous in the history of physics. Read J. D. Jackson
and L. B. Okun, “Historical roots of gauge invariance,” Rev. Mod. Phys. 73, 2001 and learn about the sad story of
a great physicist whose misfortune in life was that his name differed from that of another physicist by only one
letter.
II.7. Proof of Gauge Invariance | 145
You will notice that in (3) I have carefully given two equivalent forms. While it is simpler,
and commonly done in most textbooks, to write the second form, we should also keep thefirst form in mind. Note that /Lambda1(x) and/Lambda1(x)+2πgive exactly the same transformation.
Mathematically speaking, the quantities e
i/Lambda1(x)and∂μ/Lambda1(x) are well defined, but /Lambda1(x) is
not.
After these apparently formal but actually physically important remarks, we are ready
to work on the proof. I will let you give the general proof, but I will show you the way byworking through some representative examples.
Recall that the propagator for the hypothetical massive photon is iD
μν=i(kμkν/μ2−
gμν)/(k2−μ2). We can set the μ2in the denominator equal to zero without further ado
and write the photon propagator effectively as iDμν=i(kμkν/μ2−gμν)/k2. The dangerous
term is kμkν/μ2. We want to show that it goes away.
A specific example
First consider electron-electron scattering to order e4. Of the many diagrams, focus on the
two in figure II.7.1.a. The Feynman amplitude is then
¯u(p/prime)/parenleftbigg
γλ 1
/negationslashp+ /negationslashk−mγμ+γμ 1
/negationslashp/prime− /negationslashk−mγλ/parenrightbigg
u(p)i
k2/parenleftbiggkμkν
μ2−δν
μ/parenrightbigg
/Gamma1λν (5)
where /Gamma1λνis some factor whose detailed structure does not concern us. For the specific
case shown in figure II.7.1a we can of course write out /Gamma1λνexplicitly if we want. Note the
plus sign here from interchanging the two photons since photons obey Bose statistics.
p + kλ
μp’ − k
λμ
(a)p’
k
pp’
k
pk’
k’
Figure II.7.1
146 | II. Dirac and the Spinor
(b)p’
k
pp’
k
pp’ − k
λμ
p + kλ
μ
Figure II.7.1 (continued )
Focus on the dangerous term. Contracting the ¯u(p/prime)(...)u(p) factor in (5) with kμwe
have
¯u(p/prime)/parenleftbigg
γλ 1
/negationslashp+ /negationslashk−m/negationslashk+ /negationslashk1
/negationslashp/prime− /negationslashk−mγλ/parenrightbigg
u(p) (6)
The trick is to write the /negationslashkin the numerator of the first term as (/negationslashp+ /negationslashk−m)−(/negationslashp−m), and
in the numerator of the second term as (/negationslashp/prime−m)−(/negationslashp/prime− /negationslashk−m). Using (/negationslashp−m)u(p) =0
and¯u(p/prime)(/negationslashp/prime−m)=0, we see that the expression in (6) vanishes. This proves the theorem
in this simple example. But since the explicit form of /Gamma1λνdid not enter, the proof would
have gone through even if figure II.7.1a were replaced by the more general figure II.7.1b,where arbitrarily complicated processes could be going on under the shaded blob.
Indeed we can generalize to figure II.7.1c. Apart from the photon carrying momentum
kthat we are focusing on, there are already nphotons attached to the electron line. These
nphotons are just “spectators” in the proof in the same way that the photon carrying
momentum k
/primein figure II.7.1a never came into the proof that (6) vanishes. The photon
we are focusing on can attach to the electron line in n+1 different places. You can now
extend the proof as an exercise.
Photon landing on an internal line
In the example we just considered, the photon line in question lands on an external electron
line. The fact that the line is “capped at the two ends” by ¯u(p/prime)andu(p) is crucial in the
proof. What if the photon line in question lands on an internal line?
An example is shown in figure II.7.2, contributing to electron-electron scattering in
order e8. The figure contains three distinct diagrams. The electron “on the left” emits
II.7. Proof of Gauge Invariance | 147
(c)p’
k
p
Figure II.7.1 (continued )
three photons, which attach to an internal electron loop. The electron “on the right”
emits a photon with momentum k, which can attach to the loop in three distinct ways.
Since what we care about is whether the kμkρ/μ2piece in the photon propagator
i(kμkρ/μ2−gμρ)/k2goes away or not, we can for our purposes replace that photon
propagator by kμ. To save writing slightly, we define p1=p+q1andp2=p1+q2(see the
momentum labels in figure II.7.2): Let’s focus on the relevant part of the three diagrams,referring to them as A,B, andC.
A=/integraldisplayd4p
(2π)4tr/parenleftbigg
γν 1
/negationslashp2+ /negationslashk−mγσ 1
/negationslashp1+ /negationslashk−mγλ 1
/negationslashp+ /negationslashk−m/negationslashk1
/negationslashp−m/parenrightbigg
(7)
B=/integraldisplayd4p
(2π)4tr/parenleftbigg
γν 1
/negationslashp2+ /negationslashk−mγσ 1
/negationslashp1+ /negationslashk−m/negationslashk1
/negationslashp1−mγλ 1
/negationslashp−m/parenrightbigg
(8)
and
C=/integraldisplayd4p
(2π)4tr/parenleftbigg
γν 1
/negationslashp2+ /negationslashk−m/negationslashk1
/negationslashp2−mγσ 1
/negationslashp1−mγλ 1
/negationslashp−m/parenrightbigg
(9)
This looks like an unholy mess, but it really isn’t. We use the same trick we used before.
InCwrite/negationslashk=(/negationslashp2+ /negationslashk−m)−(/negationslashp2−m), so that
C=/integraldisplayd4p
(2π)4/bracketleftbigg
tr/parenleftbigg
γν 1
/negationslashp2−mγσ 1
/negationslashp1−mγλ 1
/negationslashp−m/parenrightbigg
−tr/parenleftbigg
γν 1
/negationslashp2+ /negationslashk−mγσ 1
/negationslashp1−mγλ 1
/negationslashp−m/parenrightbigg/bracketrightbigg
(10)
InBwrite/negationslashk=(/negationslashp1+ /negationslashk−m)−(/negationslashp1−m), so that
B=/integraldisplayd4p
(2π)4/bracketleftbigg
tr(γν 1
/negationslashp2+ /negationslashk−mγσ 1
/negationslashp1−mγλ 1
/negationslashp−m)
−tr/parenleftbigg
γν 1
/negationslashp2+ /negationslashk−mγσ 1
/negationslashp1+ /negationslashk−mγλ 1
/negationslashp−m/parenrightbigg/bracketrightbigg
(11)
148 | II. Dirac and the Spinor
p + k
k
pp1 + k
p2 + kq1
q2
−(q1 + q2 + k)λ
σ
νA
k
pp1 + k
p2 + kq1
q2
−(q1 + q2 + k)λ
σ
νBp1
kpp2
p2+kq1
q2
−(q1 + q2 + k)λ
σ
νCp1
Figure II.7.2
II.7. Proof of Gauge Invariance | 149
Finally, in Awrite/negationslashk=(/negationslashp+ /negationslashk−m)−(/negationslashp−m)
A=/integraldisplayd4p
(2π)4/bracketleftbigg
tr/parenleftbigg
γν 1
/negationslashp2+ /negationslashk−mγσ 1
/negationslashp1+ /negationslashk−mγλ 1
/negationslashp−m/parenrightbigg
−tr/parenleftbigg
γν 1
/negationslashp2+ /negationslashk−mγσ 1
/negationslashp1+ /negationslashk−mγλ 1
/negationslashp+ /negationslashk−m/parenrightbigg/bracketrightbigg
(12)
Now you see what is happening. When we add the three diagrams together terms cancel
in pairs, leaving us with
A+B+C=/integraldisplayd4p
(2π)4/bracketleftbigg
tr/parenleftbigg
γν 1
/negationslashp2−mγσ 1
/negationslashp1−mγλ 1
/negationslashp−m/parenrightbigg
−tr/parenleftbigg
γν 1
/negationslashp2+ /negationslashk−mγσ 1
/negationslashp1+ /negationslashk−mγλ 1
/negationslashp+ /negationslashk−m/parenrightbigg/bracketrightbigg
(13)
If we shift (see exercise II.7.2) the dummy integration variable p→p−kin the second
term, we see that the two terms cancel. Indeed, the kμkρ/μ2piece in the photon propagator
goes away and we can set μ=0.
I will leave the general proof to you. We have done it for one particular process. Try it
for some other process. You will see how it goes.
Ward-Takahashi identity
Let’s summarize. Given any physical amplitude Tμ...(k,...)with external electrons on
shell [this is jargon for saying that all necessary factors u(p) and¯u(p) are included in
Tμ...(k,...)] describing a process with a photon carrying momentum kcoming out of, or
going into, a vertex labeled by the Lorentz index μ, we have
kμTμ...(k,...)=0 (14)
This is sometimes known as a Ward-Takahashi identity.
The bottom line is that we can write iDμν=−igμν/k2for the photon propagator. Since
we can discard the kμkν/μ2term in the photon propagator i(kμkν/μ2−gμν)/k2we can
also add in a kμkν/k2term with an arbitrary coefficient. Thus, for the photon propagator
we can use
iDμν=i
k2/bracketleftbigg
(1−ξ)kμkν
k2−gμν/bracketrightbigg
(15)
where we can choose the number ξto simplify our calculation as much as possible.
Evidently, the choice of ξamounts to a choice of gauge for the electromagnetic field. In
particular, the choice ξ=1 is known as the Feynman gauge, and the choice ξ=0 is known
as the Landau gauge. If you find an especially nice choice, you can have a gauge named
after you as well! For fairly simple calculations, it is often advisable to calculate with anarbitrary ξ. The fact that the end result must not depend on ξprovides a useful check on
the arithmetic.
150 | II. Dirac and the Spinor
This completes the derivation of the Feynman rules for quantum electrodynamics: They
are the same rules as those given in chapter II.5 for the massive vector boson theory exceptfor the photon propagator given in (15).
We have given here a diagrammatic proof of the gauge invariance of quantum
electrodynamics. We will worry later (in chapter IV .7) about the possibility that the shift ofintegration momentum used in the proof may not be allowed in some cases.
The longitudinal mode
We now come back to the worry we had in chapter I.5. Consider a massive spin 1 mesonmoving along the z−direction. The 3 polarization vectors are fixed by the condition k
λελ=
0 with kλ=(ω,0 ,0 , k)(recall chapter I.5) and the normalization ελελ=− 1, so that ε(1)
λ=
(0, 1, 0, 0 ),ε(2)
λ=(0, 0, 1, 0 ),ε(3)
λ=(−k ,0 ,0 ,ω)/μ . Note that as μ→0, the longitudinal
polarization vector ε(3)
λbecomes proportional to kλ=(ω,0 ,0 , −k) . The amplitude for
emitting a meson with a longitudinal polarization in the process described by (14) isgiven by ε
(3)
λTλ...=(−kT0...+ωT3...)/μ=(−kT0...+/radicalbig
k2+μ2T3...)/μ/similarequal(−kT0...+
(k+μ2
2k)T3...)/μ (forμ/lessmuchk), namely −(kλTλ.../μ)+μ
2kT3...withkλ=(k,0 ,0 ,−k) . Upon
using (14) we see that the amplitude ε(3)
λTλ...→μ
2kT3...→0a sμ→0.
The longitudinal mode of the photon does not exist because it decouples from all physical
processes.
Here is an apparent paradox. Mr. Boltzmann tells us that in thermal equilibrium each
degree of freedom is associated with1
2T. Thus, by measuring some thermal property (such
as the specific heat) of a box of photon gas to an accuracy of 2 /3 an experimentalist could
tell if the photon is truly massless rather than have a mass of a zillionth of an electron volt.
The resolution is of course that as the coupling of the longitudinal mode vanishes as
μ→0 the time it takes for the longitudinal mode to come to thermal equilibrium goes to
infinity. Our crafty experimentalist would have to be very patient.
Emission and absorption of photons
According to chapter II.5, the amplitude for emitting or absorbing an external on-shell
photon with momentum kand polarization a( a=1, 2)is given by ε(a)
μ(k)Tμ...(k,...).
Thanks to (14), we are free to vary the polarization vector
ε(a)
μ(k)→ε(a)
μ(k)+λkμ (16)
for arbitrary λ. You should recognize (16) as the momentum space version of (3). As we
will see in the next chapter, by a judicious choice of ε(a)
μ(k), we can simplify a given
calculation considerably. In one choice, known as the “transverse gauge,” the 4-vectorsε
(a)
μ(k)=(0,/vectorε(k)) fora=1, 2 do not have time components. (For a photon moving in the
z-direction, this is just the choice specified in the preceding section.)
II.7. Proof of Gauge Invariance | 151
Exercises
II.7.1 Extend the proof to cover figure II.7.1c. [Hint: To get oriented, note that figure II.7.1b corresponds to
n=1.]
II.7.2 You might have worried whether the shift of integration variable is allowed. Rationalizing the denomi-
nators in the first integral
/integraldisplayd4p
(2π)4tr(γν 1
/negationslashp2−mγσ 1
/negationslashp1−mγλ 1
/negationslashp−m)
in (13) and imagining doing the trace, you can convince yourself that this integral is only logarithmically
divergent and hence that the shift is allowed. This issue will come up again in chapter IV .7 and we areanticipating a bit here.
II.8 Photon-Electron Scattering and Crossing
Photon scattering on an electron
We now apply what we just learned to calculating the amplitude for the Compton scattering
of a photon on an electron, namely, the process γ(k)+e(p)→γ(k/prime)+e(p/prime). First step:
draw the Feynman diagrams, and notice that there are two, as indicated in figure II.8.1.The electron can either absorb the photon carrying momentum kfirst or emit the photon
carrying momentum k
/primefirst. Think back to the spacetime stories we talked about in chapter
I.7. The plot of our biopic here is boringly simple: the electron comes along, absorbs andthen emits a photon, or emits and absorbs a photon, and then continues on its merry way.Because this is a quantum movie, the two alternate plots are shown superposed.
So, apply the Feynman rules (chapter II.5) to get (just to make the writing a bit easier,
we take the polarization vectors εandε
/primeto be real)
M=A(ε/prime,k/prime;ε,k)+(ε/prime↔ε,k/prime↔−k) (1)
where
A(ε/prime,k/prime;ε,k)=(−ie)2¯u(p/prime)/negationslashε/prime i
/negationslashp+ /negationslashk−m/negationslashεu(p)
=i(−ie)2
2pk¯u(p/prime)/negationslashε/prime(/negationslashp+ /negationslashk+m)/negationslashεu(p)(2)
In either case, absorb first or emit first, the electron is penalized for not being real, by the
factor of 1 /((p+k)2−m2)=1/(2pk) in one case, and 1 /((p−k/prime)2−m2)=− 1/(2pk/prime)in
the other.
At this point, to obtain the differential cross section, you just have take a deep breath and
calculate away. I will show you, however, that we could simplify the calculation considerably
by a clever choice of polarization vectors and of the frame of reference. (The calculation isstill a big mess, though!) For a change, we will be macho guys and not average and sumover the photon polarizations.
II.8. Photon-Electron Scattering | 153
p + k p − k’
(b) (a)p’
/H9255, k/H9255’, k ’
pp’
/H9255, k
p/H9255’, k ’
Figure II.8.1
In any case, we have εk=0 and ε/primek/prime=0. Now choose the transverse gauge introduced
in the preceding chapter, so that εandε/primehave zero time components. Then calculate in
the lab frame. Since p=(m,0 ,0 ,0 ), we have the additional relations
εp=0 (3)
and
ε/primep=0 (4)
Why is this a shrewd choice? Recall that /negationslasha/negationslashb=2ab−/negationslashb/negationslasha. Thus, we could move /negationslashppast
/negationslashεor/negationslashε/primeat the rather small cost of flipping a sign. Notice in (1) that (/negationslashp+ /negationslashk+m)/negationslashεu(p) =
/negationslashε(−/negationslashp− /negationslashk+m)u(p) =−/negationslashε/negationslashku(p) (where we have used εk=0.) Thus
A(ε/prime,k/prime;ε,k)=ie2¯u(p/prime)/negationslashε/prime/negationslashε/negationslashk
2pku(p) (5)
To obtain the differential cross section, we need |M|2. We will wimp out a bit and
suppose, just as in chapter II.6, that the initial electron is unpolarized and the polarizationof the final electron is not measured. Then averaging over initial polarization and summingover final polarization we have [applying II.2.8]
1
2/Sigma1/Sigma1|A(ε/prime,k/prime;ε,k)|2=e4
2(2m)2(2pk)2tr(/negationslashp/prime+m)/negationslashε/prime/negationslashε/negationslashk(/negationslashp+m)/negationslashk/negationslashε/negationslashε/prime(6)
In evaluating the trace, keep in mind that the trace of an odd number of gamma
matrices vanishes. The term proportional to m2contains /negationslashk/negationslashk=k2=0 and hence van-
ishes. We are left with tr (/negationslashp/prime/negationslashε/prime/negationslashε/negationslashk/negationslashp/negationslashk/negationslashε/negationslashε/prime)=2kptr(/negationslashp/prime/negationslashε/prime/negationslashε/negationslashk/negationslashε/negationslashε/prime)=− 2kptr(/negationslashp/prime/negationslashε/prime/negationslashε/negationslashε/negationslashk/negationslashε/prime)
=2kptr(/negationslashp/prime/negationslashε/prime/negationslashk/negationslashε/prime)=8kp[2(kε/prime)2+k/primep].
154 | II. Dirac and the Spinor
Work through the steps as indicated and the strategy should be clear. We anticommute
judiciously to exploit the “zero relations” εp=0,ε/primep=0,εk=0, and ε/primek/prime=0 and the
normalization conditions /negationslashε/negationslashε=ε2=− 1 and /negationslashε/prime/negationslashε/prime=ε/prime2=− 1 as much as possible.
We obtain
1
2/Sigma1/Sigma1|A(ε/prime,k/prime;ε,k)|2=e4
2(2m)2(2pk)28kp[2(kε/prime)2+k/primep] (7)
The other term
1
2/Sigma1/Sigma1|A(ε ,−k;ε/prime,k/prime)|2=e4
2(2m)2(2pk)28(−k/primep)[2(k/primeε)2−kp]
follows immediately by inspecting figure II.8.1 and interchanging (ε/prime↔ε,k/prime↔−k).
Just as in chapter II.6, the interference term
1
2/Sigma1/Sigma1A(ε ,−k;ε/prime,k/prime)∗A(ε/prime,k/prime;ε,k)=e4
2(2m)2(2pk)(−2pk/prime)tr(/negationslashp/prime+m)/negationslashε/prime/negationslashε/negationslashk(/negationslashp+m)/negationslashk/prime/negationslashε/prime/negationslashε (8)
is the most tedious to evaluate. Call the trace T. Clearly, it would be best to eliminate
p/prime=p+k−k/prime, since we could “do more” with /negationslashpthan/negationslashp/prime. Divide and conquer: write T=
P+Q1+Q2. First, massage P≡tr(/negationslashp+m)/negationslashε/prime/negationslashε/negationslashk(/negationslashp+m)/negationslashk/prime/negationslashε/prime/negationslashε=m2tr/negationslashε/prime/negationslashε/negationslashk/negationslashk/prime/negationslashε/prime/negationslashε+
tr/negationslashp/negationslashε/prime/negationslashε/negationslashk/negationslashp/negationslashk/prime/negationslashε/prime/negationslashε. In the second term, we could sail the first /negationslashppast the /negationslashεand/negationslashε/prime(ah, so
nice to work in the rest frame for this problem!) to find the combination /negationslashp/negationslashk/negationslashp=2kp/negationslashp−
m2/negationslashk. The m2term gives a contribution that cancels the first term in P, leaving us with
P=2kptr/negationslashε/prime/negationslashε/negationslashp/negationslashk/prime/negationslashε/prime/negationslashε=2kptr/negationslashp/negationslashk/prime/negationslashε/prime(2ε/primeε−/negationslashε/prime/negationslashε)/negationslashε=8(kp)(k/primep)[2(εε/prime)2−1]. Similarly,
Q1=tr/negationslashk/negationslashε/prime/negationslashε/negationslashk/negationslashp/negationslashk/prime/negationslashε/prime/negationslashε=− 2kε/primetr/negationslashk/negationslashp/negationslashk/prime/negationslashε/prime=− 8(ε/primek)2k/primepandQ2=− tr/negationslashk/prime/negationslashε/prime/negationslashε/negationslashk/negationslashp/negationslashk/prime/negationslashε/prime/negationslashε
=8(εk/prime)2kp/prime.
Putting it all together and writing kp/prime=k/primep=mω/primeandk/primep/prime=kp=mω,w ef i n d
1
2/Sigma1/Sigma1|M |2=e4
(2m)2/bracketleftbiggω/prime
ω+ω
ω/prime+4(εε/prime)2−2/bracketrightbigg
(9)
We calculate the differential cross section as in chapter II.6 with some minor differences
since we are in the lab frame, obtaining
dσ=m
(2π)22ω/bracketleftBigg/integraldisplayd3k/prime
2ω/primed3p/prime
Ep/primeδ(4)(k/prime+p/prime−k−p)/bracketrightBigg
1
2/summationdisplay/summationdisplay
|M|2(10)
As described in the appendix to chapter II.6, we could use (I.8.14) and write/integraltextd3p/prime
Ep/prime(...)=
/integraltext
d4p/primeθ(p/prime0)δ(p/prime2−m2)(...). Doing the integral over d4p/primeto knock out the 4-dimensional
delta function, we are left with a delta function enforcing the mass shell condition 0 =
p/prime2−m2=(p+k−k/prime)2−m2=2p(k−k/prime)−2kk/prime=2m(ω−ω/prime)−2ωω/prime(1−cosθ), with
θthe scattering angle of the photon. Thus, the frequency of the outgoing photon and of
the incoming photon are related by
ω/prime=ω
1+2ω
msin2θ
2(11)
giving the frequency shift that won Arthur Compton the Nobel Prize. You realize of course
that this formula, though profound at the time, is “merely” relativistic kinematics and hasnothing to do with quantum field theory per se.
II.8. Photon-Electron Scattering | 155
/H92552, k2
/H92551, k1 /H92552, k2/H92551, k1
p1 /H11002 k2p1 /H11002 k1
p2 p1p2 p1
(b) (a)
Figure II.8.2
What quantum field theory gives us is the Klein-Nishina formula (1929)
dσ
d/Omega1=1
(2m)2(e2
4π)2(ω/prime
ω)2/bracketleftbiggω/prime
ω+ω
ω/prime+4(εε/prime)2−2/bracketrightbigg
. (12)
You ought to be impressed by the year.
Electron-positron annihilation
Here and in chapter II.6 we calculated the cross sections for some interesting scattering
processes. At the end of that chapter we marvelled at the magic of theoretical physics.Even more magical is the annihilation of matter and antimatter, a process that occursonly in relativistic quantum field theory. Specifically, an electron and a positron meet andannihilate each other, giving rise to two photons: e
−(p1)+e+(p2)→γ(ε1,k1)+γ(ε2,k2).
(Annihilating into one physical, that is, on-shell, photon is kinematically impossible.)This process, often featured in science fiction, is unknown in nonrelativistic quantummechanics. Without quantum field theory, you would be clueless on how to calculate, say,the angular distribution of the outgoing photons.
But having come this far, you simply apply the Feynman rules to the diagrams in fig-
ure II.8.2, which describe the process to order e
2. We find the amplitude M=
A(k1,ε1;k2,ε2)+A(k 2,ε2;k1,ε1)(Bose statistics for the two photons!), where
A(k1,ε1;k2,ε2)=(ie)(−ie) ¯v(p 2)/negationslashε2i
/negationslashp1− /negationslashk1−m/negationslashε1u(p 1) (13)
Students of quantum field theory are sometimes confused that while the incoming electron
goes with the spinor u, the incoming positron goes with ¯v, and not with v. You could check
this by inspecting the hermitean conjugate of (II.2.10). Even simpler, note that ¯v(...)u
[with(...)a bunch of gamma matrices contracted with various momenta] transforms
correctly under the Lorentz group, while v(...)udoes not (and does not even make sense,
since they are both column spinors.) Or note that the annihilation operator dfor the
positron is associated with ¯v, notv.
156 | II. Dirac and the Spinor
I want to emphasize that the positron carries momentum p2=(+/radicalBig
/vectorp2
2+m2,/vectorp2)on its
way to that fatal rendezvous with the electron. Its energy p0
2=+/radicalBig
/vectorp2
2+m2is manifestly
positive. Nor is any physical particle traveling backward in time. The honest experimental-
ist who arranged for the positron to be produced wouldn’t have it otherwise. Remembermy rant at the end of chapter II.2?
In figure II.8.2a I have labeled the various lines with arrows indicating momentum
flow. The external particles are physical and there would have been serious legal issuesif their energies were not positive. There is no such restriction on the virtual particlebeing exchanged, though. Which way we draw the arrow on the virtual particle is purelyup to us. We could reverse the arrow, and then the momentum label would becomep
2−k2=k1−p1: the time component of this “composite” 4-vector can be either positive
or negative.
To make the point totally clear, we could also label the lines by dotted arrows showing the
flow of (electron) charge. Indeed, on the positron line, momentum and (electron) chargeflow in opposite directions.
Crossing
I now invite you to discover something interesting by staring at the expression in (13) fora while.
Got it? Does it remind you of some other amplitude?No? How about looking at the amplitude for Compton scattering in (2)?Notice that the two amplitudes could be turned into each other (up to an irrelevant sign)
by the exchange
p↔p1,k↔−k1,p/prime↔−p2,k/prime↔k2,ε↔ε1,ε/prime↔ε2,u(p)↔u(p 1),u(p/prime)↔v(p 2) (14)
This is known as crossing. Diagrammatically, we are effectively turning the diagrams in
figures II.8.1 and II.8.2 into each other by 90◦rotations. Crossing expresses in precise terms
what people who like to mumble something about negative energy traveling backward intime have in mind.
Once again, it is advantageous to work in the electron rest frame and in the transverse
gauge, so that we have ε
1p1=0 and ε2p1=0 as well as ε1k1=0 and ε2k1=0. Averaging
over the electron and positron polarizations we obtain
dσ
d/Omega1=α2
8m/parenleftbiggω1
|/vectorp|/parenrightbigg/bracketleftbiggω/prime
ω+ω
ω/prime−4(εε/prime)2+2/bracketrightbigg
(15)
withω1=m(m+E)/(m +E−pcosθ),ω2=(E−m−pcosθ)ω1/m, andp=|/vectorp|and
Ethe positron momentum and energy, respectively.
II.8. Photon-Electron Scattering | 157
xtime
y
Figure II.8.3
Special relativity and quantum mechanics require antimatter
The formalism in chapter II.2 makes it totally clear that antimatter is obligatory. For us to
be able to add the operators bandd†in (II.2.10) they must carry the same electric charge,
and thus banddcarry opposite charge. No room for argument there. Still, it would be
comforting to have a physical argument that special relativity and quantum mechanicsmandate antimatter.
Compton scattering offers a context for constructing a nice heuristic argument. Think
of the process in spacetime. We have redrawn figure II.8.1a in figure II.8.3: the electron ishit by the photon at the point x, propagates to the point y, and emits a photon. We have
assumed implicitly that (y
0−x0)>0, since we don’t know what propagating backward
in time means. (If the reader knows how to build a time machine, let me know.) Butspecial relativity tells us that another observer moving by (along the 1-direction say) wouldsee the time difference (y
/prime0−x/prime0)=coshϕ(y0−x0)−sinhϕ(y1−x1), which could be
negative for large enough boost parameter ϕ, provided that (y1−x1)>( y0−x0), that
is, if the separation between the two spacetime points xandywere spacelike. Then this
observer would see the field disturbance propagating from ytox. Since we see negative
electric charge propagating from xtoy, the other observer must see positive electric
charge propagating from ytox. Without special relativity, as in nonrelativistic quantum
mechanics, we simply write down the Sch ¨odinger equation for the electron and that is
that. Special relativity allows different observers to see different time ordering and henceopposite charges flowing toward the future.
Exercises
II.8.1 Show that averaging and summing over photon polarizations amounts to replacing the square bracket
in (9) by 2[ω/prime
ω+ω
ω/prime−sin2θ]. [Hint: We are working in the transverse gauge.]
II.8.2 Repeat the calculation of Compton scattering for circularly polarized photons.
This page intentionally left blank
Part III Renormalization and Gauge Invariance
This page intentionally left blank
III.1 Cutting Off Our Ignorance
Who is afraid of infinities? Not I, I just cut them off.
—Anonymous
An apparent sleight of hand
The pioneers of quantum field theory were enormously puzzled by the divergent integralsthat they often encountered in their calculations, and they spent much of the 1930sand 1940s struggling with these infinities. Many leading lights of the day, driven todesperation, advocated abandoning quantum field theory altogether. Eventually, a so-calledrenormalization procedure was developed whereby the infinities were argued away andfinite physical results were obtained. But for many years, well into the late 1960s and eventhe 1970s many physicists looked upon renormalization theory suspiciously as a sleight ofhand. Jokes circulated that in quantum field theory infinity is equal to zero and that underthe rug in a field theorist’s office had been swept many infinities.
Eventually, starting in the 1970s a better understanding of quantum field theory was
developed through the efforts of Ken Wilson and many others. Field theorists graduallycame to realize that there is no problem of divergences in quantum field theory at all. Wenow understand quantum field theory as an effective low energy theory in a sense I willexplain briefly here and in more detail in chapter VIII.3.
Field theory blowing up
We have to see an infinity before we can talk about how to deal with infinities. Well, we
saw one in chapter I.7. Recall that the order λ2correction (I.7.23) to the meson-meson
scattering amplitude diverges. With K≡k1+k2, we have
M=1
2(−iλ)2i2/integraldisplayd4k
(2π)41
k2−m2+iε1
(K−k)2−m2+iε(1)
As I remarked back in chapter I.7, even without doing any calculations we can see the
problem that confounded the pioneers of quantum field theory. The integrand goes as 1 /k4
162 | III. Renormalization and Gauge Invariance
for large kand thus the integral diverges logarithmically as/integraltext
d4k/k4. (The ordinary integral/integraltext∞dr rndiverges linearly for n=0, quadratically for n=1, and so on, and/integraltext∞dr/r
diverges logarithmically.) Since this divergence is associated with large values of kit is
known as an ultraviolet divergence.
To see how to deal with this apparent infinity, we have to distinguish between two con-
ceptually separate issues, associated with the terrible names “regularization” and “renor-malization” for historical reasons.
Parametrization of ignorance
Suppose we are studying quantum electrodynamics instead of this artificial ϕ4theory. It
would be utterly unreasonable to insist that the theory of an electron interacting with aphoton would hold to arbitrarily high energies. At the very least, with increasingly higherenergies other particles come in, and eventually electrodynamics becomes merely part ofa larger electroweak theory. Indeed, these days it is thought that as we go to higher andhigher energies the whole edifice of quantum field theory will ultimately turn out to be anapproximation to a theory whose identity we don’t yet know, but probably a string theoryaccording to some physicists.
The modern view is that quantum field theory should be regarded as an effective low
energy theory, valid up to some energy (or momentum in a Lorentz invariant theory) scale/Lambda1. We can imagine living in a universe described by our toy ϕ
4theory. As physicists in
this universe explore physics to higher and higher momentum scales they will eventuallydiscover that their universe is a mattress constructed out of mass points and springs. Thescale/Lambda1is roughly the inverse of the lattice spacing.
When I teach quantum field theory, I like to write “Ignorance is no shame” on the
blackboard for emphasis when I get to this point. Every physical theory should have adomain of validity beyond which we are ignorant of the physics. Indeed were this not truephysics would not have been able to progress. It is a good thing that Feynman, Schwinger,Tomonaga, and others who developed quantum electrodynamics did not have to knowabout the charm quark for example.
I emphasize that /Lambda1should be thought of as physical, parametrizing our threshold of
ignorance, and not as a mathematical construct.
1Indeed, physically sensible quantum
field theories should all come with an implicit /Lambda1. If anyone tries to sell you a field theory
claiming that it holds up to arbitrarily high energies, you should check to see if he soldused cars for a living. (As I wrote this, a colleague who is an editor of Physical Review Letters
told me that he worked as a garbage collector during high school vacations, adding jokinglythat this experience prepared him well for his present position.)
1We saw a particularly vivid example of this in chapter I.8. When we define a conducting plate as a surface
on which a tangential electric field vanishes, we are ignorant of the physics of the electrons rushing about tocounter any such imposed field. At extremely high frequencies, the electrons can’t rush about fast enough andnew physics comes in, namely that high frequency modes do not see the plates. In calculating the Casimir force
we parametrize our ignorance with a∼/Lambda1
−1.
III.1. Cutting Off Our Ignorance | 163
Figure III.1.1
Thus, in evaluating (1) we should integrate only up to /Lambda1, known as a cutoff. We
literally cut off the momentum integration (fig. III.1.1).2The integral is said to have been
“regularized.”
Since my philosophy in this book is to emphasize the conceptual rather than the
computational, I will not actually do the integral but merely note that it is equal to2iClog(/Lambda1
2/K2)where Cis some numerical constant that you can compute if you want
(see appendix 1 to this chapter). For the sake of simplicity I also assumed that m2<< K2
so that we could neglect m2in the integrand. It is convenient to use the kinematic variables
s≡K2=(k1+k2)2,t≡(k1−k3)2, andu≡(k1−k4)2introduced in chapter II.6. (Writing
out the kj’s explicitly in the center-of-mass frame, you see that s,t, anduare related to
rather mundane quantities such as the center-of-mass energy and the scattering angle.)After all this, the meson-meson scattering amplitude reads
M=−iλ+iCλ2[log/parenleftbigg/Lambda12
s/parenrightbigg
+log/parenleftbigg/Lambda12
t/parenrightbigg
+log/parenleftbigg/Lambda12
u/parenrightbigg
]+O(λ3) (2)
This much is easy enough to understand. After regularization, we speak of cutoff-
dependent quantities instead of divergent quantities, and Mdepends logarithmically on
the cutoff.
2A. Zee, Einstein ’s Universe , p. 204. Cartooning schools apparently teach that physicists in general, and
quantum field theorists in particular, all wear lab coats.
164 | III. Renormalization and Gauge Invariance
What is actually measured
Now that we have dealt with regularization, let us turn to renormalization, a terrible word
because it somehow implies we are doing normalization again when in fact we haven’tyet.
The key here is to imagine what we would tell an experimentalist about to measure
meson-meson scattering. We tell her (or him if you insist) that we need a cutoff /Lambda1and she
is not bothered at all; to an experimentalist it makes perfect sense that any given theoryhas a finite domain of validity.
Our calculation is supposed to tell her how the scattering will depend on the center-of-
mass energy and the scattering angle. So we show her the expression in (2). She points toλand exclaims, “What in the world is that?”
We answer, “The coupling constant,” but she says, “What do you mean, coupling
constant, it’s just a Greek letter!”
A confused student, Confusio, who has been listening in, pipes up, “Why the fuss? I
have been studying physics for years and years, and the teachers have shown us lots ofequations with Latin and Greek letters, for example, Hooke’s law F=−kx, and nobody
gets upset about kbeing just a Latin letter.”
Smart Experimentalist: “But that is because if you give me a spring I can go out and
measure k. That’s the whole point! Mr. Egghead Theorist here has to tell me how to
measure this λ.”
Woah, that is a darn smart experimentalist. We now have to think more carefully what
a coupling constant really means. Think about α, the coupling constant of quantum elec-
trodynamics. Well, it is the coefficient of 1 /rin Coulomb’s law. Fine, Monsieur Coulomb
measured it using metallic balls or something. But a modern experimentalist could justas well have measured αby scattering an electron at such and such an energy and at
such and such a scattering angle off a proton. We explain all this to our experimentalistfriend.
SE, nodding, agrees: “Oh yes, recently my colleague so and so measured the coupling for
meson-meson interaction by scattering one meson off another at such and such an energyand at such and such a scattering angle, which correspond to your variables s,t, andu
having values s
0,t0, andu0. But what does the coupling constant my colleague measured,
let us call it λP, with the subscript meaning “physical,” have to do with your theoretical λ,
which, as far as I am concerned, is just a Greek letter in something you call a Lagrangian!”
Confusio, “Hey, if she’s going to worry about small lambda, I am going to worry about
big lambda. How do I know how big the domain of validity is?”
SE: “Confusio, you are not as dumb as you look! Mr. Egghead Theorist, if I use your
formula (2), what is the precise value of /Lambda1that I am supposed to plug in? Does it depend
on your mood, Mr. Theorist? If you wake up feeling optimistic, do you use 2 /Lambda1instead of
/Lambda1? And if your girl friend left you, you use1
2/Lambda1?”
We assert, “Ha, we know the answer to that one. Look at (2): Mis supposed to be an
actual scattering amplitude and should not depend on /Lambda1. If someone wants to change /Lambda1
III.1. Cutting Off Our Ignorance | 165
we just shift λin such a way so that Mdoes not change. In fact, a couple lines of arithmetic
will show you precisely what dλ/d/Lambda1 has to be (see exercise III.1.3).”
SE: “Okay, so λis secretly a function of /Lambda1. Your notation is lousy.”
We admit, “Exactly, this bad notation has confused generations of physicists.”SE: “I am still waiting to hear how the λ
Pmy experimental colleague measured is related
to your λ.”
We say, “Aha, that’s easy. Just look at (2), which is repeated here for clarity and for your
reading convenience:
M=−iλ+iCλ2/bracketleftbigg
log/parenleftbigg/Lambda12
s/parenrightbigg
+log/parenleftbigg/Lambda12
t/parenrightbigg
+log/parenleftbigg/Lambda12
u/parenrightbigg/bracketrightbigg
+O(λ3) (3)
According to our theory, λPis given by
−iλP=−iλ+iCλ2/bracketleftbigg
log/parenleftbigg/Lambda12
s0/parenrightbigg
+log/parenleftbigg/Lambda12
t0/parenrightbigg
+log/parenleftbigg/Lambda12
u0/parenrightbigg/bracketrightbigg
+O(λ3) (4)
To show you clearly what is involved, let us denote the sum of logarithms in the square
bracket in (3) and in (4) by Land by L0, respectively, so that we can write (3) and (4) more
compactly as
M=−iλ+iCλ2L+O(λ3) (5)
and
−iλP=−iλ+iCλ2L0+O(λ3) (6)
That is how λPandλare related.”
SE: “If you give me the scattering amplitude expressed in terms of the physical coupling
λPthen it’s of use to me, but it’s not of use in terms of λ. I understand what λPis, but
notλ.”
We answer: “Fine, it just takes two lines of algebra to eliminate λin favor of λP. Big deal.
Solving (6) for λgives
−iλ=−iλP−iCλ2L0+O(λ3)=−iλP−iCλ2
PL0+O(λ3
P) (7)
The second equality is allowed to the order of approximation indicated. Now plug this
into (5)
M=−iλ+iCλ2L+O(λ3)=−iλP−iCλ2
PL0+iCλ2
PL+O(λ3
P) (8)
Please check that all manipulations are legitimate up to the order of approximation indi-
cated.”
The “miracle”
Lo and behold! The miracle of renormalization!
Now in the scattering amplitude Mwe have the combination L−L0=[log(s0/s)+
log(t0/t)+log(u0/u)]. In other words, the scattering amplitude comes out as
M=−iλP+iCλ2
P/bracketleftbigg
log/parenleftbiggs0
s/parenrightbigg
+log/parenleftbiggt0
t/parenrightbigg
+log/parenleftbiggu0
u/parenrightbigg/bracketrightbigg
+O(λ3
P) (9)
166 | III. Renormalization and Gauge Invariance
We announce triumphantly to our experimentalist friend that when the scattering
amplitude is expressed in terms of the physical coupling constant λPas she had wanted,
the cutoff /Lambda1disappear completely!
The answer should always be in terms of physically measurable quantities
The lesson here is that we should express physical quantities not in terms of “fictitious”
theoretical quantities such as λ, but in terms of physically measurable quantities such as
λP.
By the way, in the literature, λPis often denoted by λRand for historical reasons called the
“renormalized coupling constant.” I think that the physics of “renormalization” is muchclearer with the alternative term “physical coupling constant,” hence the subscript P.W e
never did have a “normalized coupling constant.”
Suddenly Confusio pipes up again; we have almost forgotten him!Confusio: “You started out with an Min (2) with two unphysical quantities λand/Lambda1,
and their “unphysicalness” sort of cancel each other out.”
SE: “Yeah, it is reminiscent of what distinguishes the good theorists from the bad ones.
The good ones always make an even number of sign errors, and the bad ones always makean odd number.”
Integrating over only the slow modes
In the path integral formulation, the scattering amplitude Mdiscussed here is obtained
by evaluating the integral (chapter I.7)
/integraldisplay
Dϕ ϕ(x 1)ϕ(x 2)ϕ(x 3)ϕ(x4)ei/integraltext
ddx{1
2[(∂ϕ)2−m2ϕ2]−λ
4!ϕ4}
The regularization used here corresponds roughly to restricting ourselves, in the integral/integraltext
Dϕ, to integrating over only those field configurations ϕ(x) whose Fourier transform
ϕ(k) vanishes for k>∼/Lambda1. In other words, the fields corresponding to the internal lines in
the Feynman diagrams in fig. (I.7.10) are not allowed to fluctuate too energetically. We willcome back to this path integral formulation later when we discuss the renormalizationgroup.
Alternative lifestyles
I might also mention that there are a number of alternative ways of regularizing Feyn-
man diagrams, each with advantages and disadvantages that make them suitable for somecalculations but not others. The regularization used here, known as Pauli-Villars, has theadvantage of being physically transparent. Another often used regularization is known as
III.1. Cutting Off Our Ignorance | 167
dimensional regularization. We pretend that we are calculating in d-dimensional space-
time. After the Feynman integral has been beaten down to a suitable form, we do an analyticcontinuation in dand set d=4 at the end of the day. The cutoff dependences of various
integrals now show up as poles as we let d→4. Just as the cutoff /Lambda1disappears when
the scattering amplitude is expressed in terms of the physical coupling constant λ
P, in di-
mensional regularization the scattering amplitude expressed in terms of λPis free of poles.
While dimensional regularization proves to be useful in certain contexts as I will note in alater chapter, it is considerably more abstract and formal than Pauli-Villars regularization.Each to his or her own taste when it comes to regularizing.
Since the emphasis in this book is on the conceptual rather than the computational, I
won’t discuss other regularization schemes but will merely sketch how Pauli-Villars anddimensional regularizations work in two appendixes to this chapter.
Appendix 1: Pauli-Villars regularization
The important message of this chapter is the conceptual point that when physical amplitudes are expressed in
term of physical coupling constants the cutoff dependence disappears. The actual calculation of the Feynmanintegral is unimportant. But I will show you how to do the integral just in case you would like to do Feynmanintegrals for a living.
Let us start with the convergent integral
/integraldisplayd4k
(2π)41
(k2−c2+iε)3=−i
32π2c2(10)
The dependence on c2follows from dimensional analysis. The overall factor is calculated in appendix D.
Applying the identity (D.15)
1
xy=/integraldisplay1
0dα1
[αx+(1−α)y ]2(11)
to (1) we have
M=1
2(−iλ)2i2/integraldisplayd4k
(2π)4/integraldisplay1
0dα1
D
with
D=[α(K−k)2+(1−α)k2−m2+iε]2=[(k−αK)2+α(1−α)K2−m2+iε]2
Shift the integration variable k→k+αK and we meet the integral/integraltext
[d4k/(2π)4][ 1/(k2−c2+iε)2], where
c2=m2−α(1−α)k2. Pauli-Villars proposed replacing it by
/integraldisplayd4k
(2π)4/bracketleftbigg1
(k2−c2+iε)2−1
(k2−/Lambda12+iε)2/bracketrightbigg
(12)
with/Lambda12/greatermuchc2.F o rk much smaller than /Lambda1the added second term in the integrand is of order /Lambda1−4and is negligible
compared to the first term since /Lambda1is much larger than c.F o rkmuch larger than /Lambda1, the two terms almost cancel
and the integrand vanishes rapidly with increasing k, effectively cutting off the integral.
Upon differentiating (12) with respect to c2and using (10) we deduce that (12) must be equal to
(i/16π2)log(/Lambda12/c2). Thus, the integral
/integraldisplay/Lambda1d4k
(2π)41
(k2−c2+iε)2=i
16π2log/parenleftbigg/Lambda12
c2/parenrightbigg
(13)
is indeed logarithmically dependent on the cutoff, as anticipated in the text.
168 | III. Renormalization and Gauge Invariance
For what it is worth, we obtain
M=iλ2
32π2/integraldisplay1
0dαlog/parenleftbigg/Lambda12
m2−α(1−α)K2−iε/parenrightbigg
(14)
Appendix 2: Dimensional regularization
The basic idea behind dimensional regularization is very simple. When we reach
I=/integraltext
[d4k/(2π)4][1/(k2−c2+iε)2] we rotate to Euclidean space and generalize to ddimensions (see appen-
dix D):
I(d)=i/integraldisplaydd
Ek
(2π)d1
(k2+c2)2=i/bracketleftbigg2πd/2
/Gamma1(d/ 2)/bracketrightbigg1
(2π)d/integraldisplay∞
0dk kd−1 1
(k2+c2)2
As I said, I don’t want to get bogged down in computation in this book, but we’ve got to do what we’ve got to do.
Changing the integration variable by setting k2+c2=c2/xwe find
/integraldisplay∞
0dk kd−1 1
(k2+c2)2=1
2cd−4/integraldisplay1
0dx(1−x)d/2−1x1−d/ 2,
which we are supposed to recognize as the integral representation of the beta function. After the dust settles, we
obtain
i/integraldisplaydd
Ek
(2π)d1
(k2+c2)2=i1
(4π)d/2/Gamma1/parenleftbigg4−d
2/parenrightbigg
cd−4(15)
Asd→4, the right-hand side becomes
i1
(4π)2/bracketleftbigg2
4−d−logc2+log(4π)−γ+O(d−4)/bracketrightbigg
where γ=0.577 ...denotes the Euler-Mascheroni constant.
Comparing with (13) we see that log /Lambda12in Pauli-Villars regularization has been effectively replaced by the pole
2/(4−d). As noted in the text, when physical quantities are expressed in terms of physical coupling constants,
all such poles cancel.
Exercises
III.1.1 Work through the manipulations leading to (9) without referring to the text.
III.1.2 Regard (1) as an analytic function of K2. Show that it has a cut extending from 4 m2to infinity. [Hint: If
you can’t extract this result directly from (1) look at (14). An extensive discussion of this exercise will begiven in chapter III.8.]
III.1.3 Change /Lambda1toe
ε/Lambda1. Show that for M not to change, to the order indicated λmust change by δλ=
6εCλ2+O(λ3), that is,
/Lambda1dλ
d/Lambda1=6Cλ2+O(λ3)
III.2 Renormalizable versus Nonrenormalizable
Old view versus new view
We learned that if we were to write the meson meson scattering amplitude in terms
of a physically measured coupling constant λP, the dependence on the cutoff /Lambda1would
disappear (at least to order λ2
P). Were we lucky or what?
Well, it turns out that there are quantum field theories in which this would happen and
that there are quantum field theories in which this would not happen, which gives us abinary classification of quantum field theories. Again, for historical reasons, the formerare known as “renormalizable theories” and are considered “nice.” The latter are knownas “nonrenormalizable theories,” evoking fear and loathing in theoretical physicists.
Actually, with the new view of field theories as effective low energy theories to some
underlying theory, physicists now look upon nonrenormalizable theories in a much moresympathetic light than a generation ago. I hope to make all these remarks clear in this anda later chapter.
High school dimensional analysis
Let us begin with some high school dimensional analysis. In natural units in which /planckover2pi=1
andc=1, length and time have the same dimension, the inverse of the dimension of mass
(and of energy and momentum). Particle physicists tend to count dimension in terms ofmass as they are used to thinking of energy scales. Condensed matter physicists, on theother hand, usually speak of length scales. Thus, a given field operator has (equal and)opposite dimensions in particle physics and in condensed matter physics. We will use theconvention of the particle physicists.
Since the action S≡/integraltext
d
4xLappears in the path integral as eiS, it is clearly dimension-
less, thus implying that the Lagrangian (Lagrangian density, strictly speaking) Lhas the
same dimension as the 4th power of a mass. We will use the notation [ L]=4 to indicate
170 | III. Renormalization and Gauge Invariance
that Lhas dimension 4. In this notation [ x]=− 1 and [ ∂]=1. Consider the scalar field
theory L=1
2[(∂ϕ)2−m2ϕ2]−λϕ4. For the term (∂ϕ)2to have dimension 4, we see that
[ϕ]=1 (since 2 (1+[ϕ])=4). This then implies that [ λ]=0, that is, the coupling λis dimen-
sionless. The rule is simply that for each term in L, the dimensions of the various pieces,
including the coupling constant and mass, have to add up to 4 (thus, e.g., [ λ]+4[ϕ]=4).
How about the fermion field ψ? Applying this rule to the Lagrangian L=¯ψiγμ∂μψ+
... we see that [ ψ]=3
2. (Henceforth we will suppress the ...; it is understood that we
are looking at a piece of the Lagrangian. Furthermore, since we are doing dimensionalanalysis we will often suppress various irrelevant factors, such as numerical factors andthe gamma matrices in the Fermi interaction that we will come to presently.) Looking atthe coupling fϕ¯ψψ we see that the Yukawa coupling fis dimensionless. In contrast, in
the theory of the weak interaction with L=G¯ψψ¯ψψ we see that the Fermi coupling G
has dimension −2 (since −2+4(
3
2)=4; got that?).
From the Maxwell Lagrangian −1
4FμνFμνwe see that [ Aμ]=1 and hence Aμhas the
same dimension as ∂μ: The vector field has the same dimension as the scalar field. The
electromagnetic coupling eAμ¯ψγμψtells us that eis dimensionless, which we can also
deduce from Coulomb’s law written in natural units V( r)=α/r , with the fine structure
constant α=e2/4π.
Scattering amplitude blows up
We are now ready for a heuristic argument regarding the nonrenormalizability of a theory.
Consider Fermi’s theory of the weak interaction. Imagine calculating the amplitude M
for a four-fermion interaction, say neutrino-neutrino scattering at an energy much smallerthan/Lambda1. In lowest order, M∼G. Let us try to write down the amplitude to the next order:
M∼G+G
2(?), where we will try to guess what (?) is. Since all masses and energies
are by definition small compared to the cutoff /Lambda1, we can simply set them equal to zero.
Since [ G]=− 2, by high school dimensional analysis the unknown factor (?)must have
dimension +2. The only possibility for (?)is/Lambda12. Hence, the amplitude to the next order
must have the form M∼G+G2/Lambda12. We can also check this conclusion by looking at the
Feynman diagram in figure III.2.1: Indeed it goes as G2/integraltext/Lambda1d4p(1/p)(1/p)∼G2/Lambda12.
Without a cutoff on the theory, or equivalently with /Lambda1=∞ , theorists realized that the
theory was sick: Infinity was the predicted value for a physical quantity. Fermi’s weakinteraction theory was said to be nonrenormalizable. Furthermore, if we try to calculate tohigher order, each power of Gis accompanied by another factor of /Lambda1
2.
In desperation, some theorists advocated abandoning quantum field theory altogether.
Others expended an enormous amount of effort trying to “cure” weak interaction theory.For instance, one approach was to speculate that the series (with coefficients suppressed)
M∼G[1+G/Lambda1
2+(G/Lambda12)2+(G/Lambda12)4+...] summed to Gf (G/Lambda12), where the unknown
function fmight have the property that f(∞) was finite. In hindsight, we now know that
this is not a fruitful approach.
III.2. Renormalization Issues | 171
ν
ννν
νν
Figure III.2.1
Instead, what happened was that toward the late 1960s S. Glashow, A. Salam, and S.
Weinberg, building on the efforts of many others, succeeded in constructing an elec-troweak theory unifying the electromagnetic and weak interactions, as I will discuss inchapter VII.2. Fermi’s weak interaction theory emerges within electroweak theory as alow energy effective theory.
Fermi’s theory cried out
In modern terms, we think of the cutoff /Lambda1as really being there and we hear the cutoff
dependence of the four-fermion interaction amplitude M∼G+G2/Lambda12as the sound of the
theory crying out that something dramatic has to happen at the energy scale /Lambda1∼(1/G)1
2.
The second term in the perturbation series becomes comparable to the first, so at the veryleast perturbation theory fails.
Here is another way of making the same point. Suppose that we don’t know anything
about cutoff and all that. With Ghaving mass dimension −2, just by high school dimen-
sional analysis we see that the neutrino-neutrino scattering amplitude at center-of-mass
energy Ehas to go as M∼G+G
2E2+.... When Ereaches the scale ∼(1/G)1
2the am-
plitude reaches order unity and some new physics must take over just because the cross
section is going to violate the unitarity bound from basic quantum mechanics. (Rememberphase shift and all that?)
In fact, what that something is goes back to Yukawa, who at the same time that he
suggested the meson theory for the nuclear forces also suggested that an intermediatevector boson could account for the Fermi theory of the weak interaction. (In the 1930s thedistinction between the strong and the weak interactions was far from clear.) Schematically,consider a theory of a vector boson of mass Mcoupled to a fermion field via a dimensionless
coupling constant g:
L=¯ψ(iγμ∂μ−m)ψ−1
4FμνFμν+M2AμAμ+gAμ¯ψγμψ (1)
172 | III. Renormalization and Gauge Invariance
k
Figure III.2.2
Let’s calculate fermion-fermion scattering. The Feynman diagram in figure III.2.2 gen-
erates an amplitude (−ig)2(¯uγμu)[i/(k2−M2+iε)](¯uγμu), which when the momentum
transfer kis much less than Mbecomes i(g2/M2)(¯uγμu)(¯uγμu). But this is just as if the
fermions are interacting via a Fermi theory of the form G(¯ψγμψ)(¯ψγμψ)withG=g2/M2.
If we blithely calculate with the low energy effective theory G(¯ψγμψ)(¯ψγμψ), it cries
out that it is going to fail. Yes sir indeed, at the energy scale (1/G)1
2=M/g , the vector
boson is produced. New physics appears.
I find it sobering and extremely appealing that theories in physics have the ability to
announce their own eventual failure and hence their domains of validity, in contrast totheories in some other areas of human thought.
Einstein’s theory is now crying out
The theory of gravity is also notoriously nonrenormalizable. Simply comparing Newton’slawV( r)=G
NM1M2/rwith Coulomb’s V( r)=α/r we see that Newton’s gravitational
constant GNhas mass dimension −2. No more need be said. We come to the same
morose conclusion that the theory of gravity, just like Fermi’s theory of weak interaction,is nonrenormalizable. To repeat the argument, if we calculate graviton-graviton scatteringat energy E, we encounter the series ∼[1+G
NE2+(GNE2)2+...].
Just as in our discussion of the Fermi theory, the nonrenormalizability of quantum grav-
ity tells us that at the Planck energy scale (1/GN)1
2≡MPlanck∼1019mproton new physics
must appear. Fermi’s theory cried out, and the new physics turned out to be the elec-
troweak theory. Einstein’s theory is now crying out. Will the new physics turn out to bestring theory?
1
Exercise
III.2.1 Consider the d-dimensional scalar field theory S=/integraltext
ddx(1
2(∂ϕ)2+1
2m2ϕ2+λϕ4+...+λnϕn+...).
Show that [ ϕ]=(d−2)/2 and [ λn]=n(2−d)/2+d. Note that ϕis dimensionless for d=2.
1J. Polchinski, String Theory .
III.3 Counterterms and Physical Perturbation Theory
Renormalizability
The heuristic argument of the previous chapter indicates that theories whose coupling has
negative mass dimension are nonrenormalizable. What about theories with dimensionlesscouplings, such as quantum electrodynamics and the ϕ
4theory? As a matter of fact, both
of these theories have been proved to be renormalizable. But it is much more difficult toprove that a theory is renormalizable than to prove that it is nonrenormalizable. Indeed,the proof that nonabelian gauge theory (about which more later) is renormalizable tookthe efforts of many eminent physicists, culminating in the work of ’t Hooft, Veltman, B.Lee, Zinn-Justin, and many others.
Consider again the simple ϕ
4theory. First, a trivial remark: The physical coupling
constant λPis a function of s0,t0, andu0[see (III.1.4)]. For theoretical purposes it is much
less cumbersome to set s0,t0, and u0equal to μ2and thus use, instead of (III.1.4), the
simpler definition
−iλP=−iλ+3iCλ2log/parenleftbigg/Lambda12
μ2/parenrightbigg
+O(λ3) (1)
This is purely for theoretical convenience.1
We saw that to order λ2the meson-meson scattering amplitude when expressed in terms
of the physical coupling λPis independent of the cutoff /Lambda1. How do we prove that this is true
to all orders in λ? Dimensional analysis only tells us that to any order in λthe dependence of
the meson scattering amplitude on the cutoff must be a sum of terms going as [log (/Lambda1/μ)]p
with some power p.
The meson-meson scattering amplitude is certainly not the only quantity that depends
on the cutoff. Consider the inverse of the ϕpropagator to order λ2as shown in figure III.3.1.
1In fact, the kinematic point s0=t0=u0=μ2cannot be reached experimentally, but that’s of concern to
theorists.
174 | III. Renormalization and Gauge Invariance
kkq
kk
qp
(a) (b)
Figure III.3.1
The Feynman diagram in figure III.3.1a gives something like
−iλ/integraldisplay/Lambda1/bracketleftbiggd4q
(2π)4/bracketrightbigg/bracketleftbiggi
q2−m2+iε/bracketrightbigg
The precise value does not concern us; we merely note that it depends quadratically on the
cutoff /Lambda1but not on k2. The diagram in figure III.3.1b involves a double integral
I(k,m,/Lambda1;λ)≡(−iλ)2/integraldisplay/Lambda1/integraldisplay/Lambda1d4p
(2π)4d4q
(2π)4
i
p2−m2+iεi
q2−m2+iεi
(p+q+k)2−m2+iε(2)
Counting powers of pandqwe see that the integral ∼/integraltext
(d8P/P6)and so Idepends
quadratically on the cutoff /Lambda1.
By Lorentz invariance Iis a function of k2, which we can expand in a series D+Ek2+
Fk4+.... The quantity Dis just Iwith the external momentum kset equal to zero and
so depends quadratically on the cutoff /Lambda1. Next, we can obtain Eby differentiating Iwith
respect to ktwice and then setting kequal to zero. This clearly decreases the powers of pand
qin the integrand by 2 and so Edepends only logarithmically on the cutoff /Lambda1. Similarly,
we can obtain Fby differentiating Iwith respect to kfour times and then setting kequal to
zero. This decreases the powers of pandqin the integrand by 4 and thus Fis given by an
integral that goes as ∼/integraltext
d8P/P10for large P. The integral is convergent and hence cutoff
independent. We can clearly repeat the argument ad infinitum. Thus Fand the terms in
(...)are cutoff independent as the cutoff goes to infinity and we don’t have to worry about
them.
Putting it altogether, we have the inverse propagator k2−m2+a+bk2up toO(k2)with
aandb, respectively, quadratically and logarithmically cutoff dependent. The propagator
is changed to
1
k2−m2→1
(1+b)k2−(m2−a)(3)
The pole in k2is shifted to m2
P≡m2+δm2≡(m2−a)(1+b)−1, which we identify as
the physical mass. This shift is known as mass renormalization. Physically, it is quitereasonable that quantum fluctuations will shift the mass.
III.3. Physical Perturbation Theory | 175
What about the fact that the residue of the pole in the propagator is no longer 1 but
(1+b)−1?
To understand this shift in the residue, recall that we blithely normalized the field ϕso
that L=1
2(∂ϕ)2+.... That the coefficient of k2in the lowest order inverse propagator
k2−m2is equal to 1 reflects the fact that the coefficient of1
2(∂ϕ)2inLis equal to 1.
There is certainly no guarantee that with higher order corrections included the coefficientof
1
2(∂ϕ)2in an effective Lwill stay at 1. Indeed, we see that it is shifted to (1+b).F o r
historical reasons, this is known as “wave function renormalization” even though there isno wave function anywhere in sight. A more modern term would be field renormalization.(The word renormalization makes some sense in this case, as we did normalize the fieldwithout thinking too much about it.)
Incidentally, it is much easier to say “logarithmic divergent” than to say “logarithmically
dependent on the cutoff /Lambda1,” so we will often slip into this more historical and less accurate
jargon and use the word divergent. In ϕ
4theory, the wave function renormalization and the
coupling renormalization are logarithmically divergent, while the mass renormalizationis quadratically divergent.
Bare versus physical perturbation theory
What we have been doing thus far is known as bare perturbation theory. We should haveput the subscript 0 on what we have been calling ϕ,m, andλ. The field ϕ
0is known as
the bare field, and m0andλ0are known as the bare mass and bare coupling, respectively.
I did not put on the subscript 0 way back in part I because I did not want to clutter up thenotation before you, the student, even knew what a field was.
Seen in this light, using bare perturbation theory seems like a really stupid thing to
do, and it is. Shouldn’t we start out with a zeroth order theory already written in terms ofthe physical mass m
Pand physical coupling λPthat experimentalists actually measure,
and perturb around that theory? Yes, indeed, and this way of calculating is known asrenormalized or dressed perturbation theory, or as I prefer to call it, physical perturbationtheory.
We write
L=1
2[(∂ϕ)2−m2
Pϕ2]−λP
4!ϕ4+A(∂ϕ)2+Bϕ2+Cϕ4(4)
(A word on notation: The pedantic would probably want to put a subscript Pon the field
ϕ, but let us clutter up the notation as little as possible.) Physical perturbation theory
works as follows. The Feynman rules are as before, but with the crucial difference thatfor the coupling we use λ
Pand for the propagator we write i/(k2−m2
P+iε) with the
physical mass already in place. The last three terms in (4) are known as counterterms.
The coefficients A,B, andCare determined iteratively (see later) as we go to higher and
higher order in perturbation theory. They are represented as crosses in Feynman diagrams,as indicated in figure III.3.2, with the corresponding Feynman rules. All momentumintegrals are cut off.
176 | III. Renormalization and Gauge Invariance
k+2i(Ak2 + B)i
k2−mp2k
k−iλp
4!iC
Figure III.3.2
Let me now explain how A,B, and Care determined iteratively. Suppose we have
determined them to order λN
P. Call their values to this order AN,BN,andCN. Draw all the
diagrams that appear in order λN+1
P. We determine AN+1,BN+1, andCN+1by requiring
that the propagator calculated to the order λN+1
Phas a pole at mPwith a residue equal to
1, and that the meson-meson scattering amplitude evaluated at some specified values ofthe kinematic variables has the value −iλ
P. In other words, the counterterms are fixed
by the condition that mPandλPare what we say they are. Of course, A,B, andCwill
be cutoff dependent. Note that there are precisely three conditions to determine the threeunknowns A
N+1,BN+1, andCN+1.
Explained in this way, you can see that it is almost obvious that physical perturbation
theory works, that is, it works in the sense that all the physical quantities that we calculatewill be cutoff independent. Imagine, for example, that you labor long and hard to calculatethe meson-meson scattering amplitude to order λ
17
P. It would contain some cutoff depen-
dent and some cutoff independent terms. Then you simply add a contribution given byC
17and adjust C17to cancel the cutoff dependent terms.
But ah, you start to worry. You say, “What if I calculate the amplitude for two mesons to
go into four mesons, that is, diagrams with six external legs? If I get a cutoff dependentanswer, then I am up the creek, as there is no counterterm of the form Dϕ
6in (4) to soak
up the cutoff dependence.” Very astute of you, but this worry is covered by the following
power counting theorem.
Degree of divergence
Consider a diagram with BEexternal ϕlines. First, a definition: A diagram is said to have
a superficial degree of divergence Dif it diverges as /Lambda1D. (A logarithmic divergence log /Lambda1
counts as D=0.)The theorem says that Dis given by
D=4−BE (5)
III.3. Physical Perturbation Theory | 177
Figure III.3.3
I will give a proof later, but I will first illustrate what (5) means. For the inverse propagator,
which has BE=2, we are told that D=2. Indeed, we encountered a quadratic divergence.
For the meson-meson scattering amplitude, BE=4, and so D=0, and indeed, it is
logarithmically divergent.
According to the theorem, if you calculate a diagram with six external legs (what is
technically sometimes known as the six-point function), BE=6 andD=− 2. The theorem
says that your diagram is convergent or cutoff independent (i.e., the cutoff dependencedisappears as /Lambda1
−2). You didn’t have to worry. You should draw a few diagrams to check
this point. Diagrams with more external legs are even more convergent.
The proof of the theorem follows from simple power counting. In addition to BEand
D, let us define BIas the number of internal lines, Vas the number of vertices, and L
as the number of loops. (It is helpful to focus on a specific diagram such as the one infigure III.3.3 with B
E=6,D=− 2,BI=5,V=4, and L=2.)
The number of loops is just the number of/integraltext
[d4k/(2π)4] we have to do. Each internal
line carries with it a momentum to be integrated over, so we seem to have BIintegrals to
do. But the actual number of integrals to be done is of course decreased by the momentumconservation delta functions associated with the vertices, one to each vertex. There are thusVdelta functions, but one of them is associated with overall momentum conservation of
the entire diagram. Thus, the number of loops is
L=BI−(V−1) (6)
[If you have trouble following this argument, you should work things out for the diagram
in figure III.3.3 for which this equation reads 2 =5−(4−1).]
For each vertex there are four lines coming out (or going in, depending on how you look
at it). Each external line comes out of (or goes into) one vertex. Each internal line connectstwo vertices. Thus,
4V=BE+2BI (7)
(For figure III.3.3 this reads 4 .4=6+2.5.)
178 | III. Renormalization and Gauge Invariance
Finally, for each loop there is a/integraltext
d4kwhile for each internal line there is a
i/(k2−m2+iε), bringing the powers of momentum down by 2. Hence,
D=4L−2BI (8)
(For figure III.3.3 this reads −2=4.2−2.5.)
Putting (6), (7), and (8) together, we obtain the theorem (5).2As you can plainly see, this
formalizes the power counting we have been doing all along.
Degree of divergence with fermions
To test if you understood the reasoning leading to (5), consider the Yukawa theory we metin (II.5.19). (We suppress the counterterms for typographical clarity.)
L=¯ψ(iγμ∂μ−mP)ψ+1
2[(∂ϕ)2−μ2
Pϕ2]−λPϕ4+fPϕ¯ψψ (9)
Now we have to count FIandFE, the number of internal and external fermion lines,
respectively, and keep track of VfandVλ, the number of vertices with the coupling fand
λ, respectively. We have five equations altogether. For instance, (7) splits into two equationsbecause we now have to count fermion lines as well as boson lines. For example, we nowhave
Vf+4Vλ=BE+2BI (10)
We (that is, you) finally obtain
D=4−BE−3
2FE (11)
So the divergent amplitudes, that is, those classes of diagrams with D≥0, have (BE,FE)=
(0, 2),(2, 0),(1, 2), and(4, 0). We see that these correspond to precisely the six terms in
the Lagrangian (9), and thus we need six counterterms.
Note that this counting of superficial powers of divergence shows that all terms with
mass dimension ≤4 are generated. For example, suppose that in writing down the La-
grangian (9) we forgot to include the λPϕ4term. The theory would demand that we include
this term: We have to introduce it as a counterterm [the term with (BE,FE)=(4, 0)in the
list above].
A common feature of (5) and (11) is that they both depend only on the number of external
lines and not on the number of vertices V. Thus, for a given number of external lines, no
matter to what order of perturbation theory we go, the superficial degree of divergenceremains the same. Further thought reveals that we are merely formalizing the dimension-counting argument of the preceding chapter. [Recall that the mass dimension of a Bosefield [ϕ] is 1 and of a Fermi field [ ψ]i s
3
2. Hence the coefficients 1 and3
2in (11).]
2The superficial degree of divergence measures the divergence of the Feynman diagram as all internal
momenta are scaled uniformly by k→akwithatending to infinity. In a more rigorous treatment we have to
worry about the momenta in some subdiagram (a piece of the full diagram) going to infinity with other momenta
held fixed.
III.3. Physical Perturbation Theory | 179
Our discussion hardly amounts to a rigorous proof that theories such as the Yukawa
theory are renormalizable. If you demand rigor, you should consult the many field theorytomes on renormalization theory, and I do mean tomes, in which such arcane topics asoverlapping divergences and Zimmerman’s forest formula are discussed in exhaustive andexhausting detail.
Nonrenormalizable field theories
It is instructive to see how nonrenormalizable theories reveal their unpleasant personali-ties when viewed in the context of this discussion. Consider the Fermi theory of the weakinteraction written in a simplified form:
L=¯ψ(iγμ∂μ−mP)ψ+G(¯ψψ)2.
The analogs of (6), (7), and (8) now read L=FI−(V−1),4V=FE+2FI, and D=
4L−FI. Solving for the superficial degree of divergence in terms of the external number
of fermion lines, we find
D=4−3
2FE+2V (12)
Compared to the corresponding equations for renormalizable theories (5) and (11), D
now depends on V. Thus if we calculate fermion-fermion scattering (FE=4), for example,
the divergence gets worse and worse as we go to higher and higher order in the perturbationseries. This confirms the discussion of the previous chapter. But the really bad news isthat for any F
E, we would start running into divergent diagrams when Vgets sufficiently
large, so we would have to include an unending stream of counterterms (¯ψψ)3,(¯ψψ)4,
(¯ψψ)5,..., each with an arbitrary coupling constant to be determined by an experimental
measurement. The theory is severely limited in predictive power.
At one time nonrenormalizable theories were considered hopeless, but they are accepted
in the modern view based on the effective field theory approach, which I will discuss inchapter VIII.3.
Dependence on dimension
The superficial degree of divergence clearly depends on the dimension dof spacetime since
each loop is associated with/integraltext
ddk. For example, consider the Fermi interaction G(¯ψψ)2in
(1+1)-dimensional spacetime. Of the three equations that went into (12) one is changed
toD=2L−FI, giving
D=2−1
2FE (13)
the analog of (12) for 2-dimensional spacetime. In contrast to (12), Vno longer enters, and
the only superficial diagrams have FE=2 and 4, which we can cancel by the appropriate
180 | III. Renormalization and Gauge Invariance
counterterms. The Fermi interaction is renormalizable in (1+1)-dimensional spacetime.
I will come back to it in chapter VII.4.
The Weisskopf phenomenon
I conclude by pointing out that the mass correction to a Bose field and to a Fermi fielddiverge differently. Since this phenomenon was first discovered by Weisskopf, I refer toit as the Weisskopf phenomenon. To see this go back to (11) and observe that B
EandFE
contribute differently to the superficial degree of divergence D.F o rBE=2,FE=0, we
haveD=2, and thus the mass correction to a Bose field diverges quadratically, as we have
already seen explicitly [the quantity ain (3)]. But for FE=2,BE=0, we have D=1 and
it looks like the fermion mass is linearly divergent. Actually, in 4-dimensional field theorywe cannot possibly get a linear dependence on the cutoff. To see this, it is easiest to look at,as an example, the Feynman integral you wrote down for exercise II.5.1, for the diagramin figure II.5.1:
(if )2i2/integraldisplayd4k
(2π)41
k2−μ2/negationslashp+ /negationslashk+m
(p+k)2−m2≡A(p2)/negationslashp+B(p2) (14)
where I define the two unknown functions A(p2)andB(p2)for convenience. (For the
purpose of this discussion it doesn’t matter whether we are doing bare or physical pertur-bation theory. If the latter, then I have suppressed the subscript Pfor the sake of notational
clarity.)
Look at the integrand for large k. You see that the integral goes as/integraltext
/Lambda1d4k(/negationslashk/k4)and
looks linearly divergent, but by reflection symmetry k→−kthe integral to leading order
vanishes. The integral in (14) is merely logarithmically divergent. The superficial degreeof divergence Doften gives an exaggerated estimate of how bad the divergence can be
(hence the adjective “superficial”). In fact, staring at (14) we can prove more, for instance,thatB(p
2)must be proportional to m. As an exercise you can show, using the Feynman
rules given in chapter II.5, that the same conclusion holds in quantum electrodynamics.
For a boson, quantum fluctuations give δμ∝/Lambda12/μ, while for a fermion such as the
electron, quantum correction to its mass δm∝mlog(/Lambda1/m) is much more benign. It is
interesting to note that in the early twentieth century physicists thought of the electronas a ball of charge of radius a. The electrostatic energy of such a ball, of the order e
2/a,
was identified as the electron mass. Interpreting 1 /aas/Lambda1, we could say that in classical
physics the electron mass is proportional to /Lambda1and diverges linearly. Thus, one way of
stating the Weisskopf phenomenon is that “bosons behave worse than a classical charge,but fermions behave better.”
As Weisskopf explained in 1939, the difference in the degree of the divergence can be
understood heuristically in terms of quantum statistics. The “bad” behavior of bosons has
to do with their gregariousness. A fermion would push away the virtual fermions fluctuat-ing in the vacuum, thus creating a cavity in the vacuum charge distribution surrounding
III.3. Physical Perturbation Theory | 181
it. Hence its self-energy is less singular than would be the case were quantum statistics
not taken into account. A boson does the opposite.
The “bad” behavior of bosons will come back to haunt us later.
Power of /planckover2picounts the number of loops
This is a convenient place to make a useful observation, though unrelated to divergences
and cut-off dependence. Suppose we restore Planck’s unit of action /planckover2pi. In the path integral,
the integrand becomes eiS//planckover2pi(recall chapter I.2) so that effectively L→L//planckover2pi. Consider
L=−1
2ϕ(∂2+m2)ϕ−λ
4!ϕ4, just to be definite. The coupling λ→λ//planckover2pi, so that each vertex
is now associated with a factor of 1 //planckover2pi. Recall that the propagator is essentially the inverse
of the operator (∂2+m2), and so in momentum space 1 /(k2−m2)→/planckover2pi/(k2−m2). Thus
the powers of /planckover2piare given by the number of internal lines minus the number of vertices
P=BI−V=L−1, where we used (6). You can check that this holds in general and not
just for ϕ4theory.
This observation shows that organizing Feynman diagrams by the number of loops
amounts to an expansion in Planck’s constant (sometimes called a semi-classical expan-sion), with the tree diagrams providing the leading term. We will come across this againin chapter IV .3.
Exercises
III.3.1 Show that in (1+1)-dimensional spacetime the Dirac field ψhas mass dimension1
2, and hence the
Fermi coupling is dimensionless.
III.3.2 Derive (11) and (13).
III.3.3 Show that B(p2)in (14) vanishes when we set m=0. Show that the same behavior holds in quantum
electrodynamics.
III.3.4 We showed that the specific contribution (14) to δmis logarithmically divergent. Convince yourself that
this is actually true to any finite order in perturbation theory.
III.3.5 Show that the result P=L−1 holds for all the theories we have studied.
III.4 Gauge Invariance: A Photon Can Find No Rest
When the central identity blows up
I explained in chapter I.7 that the path integral for a generic field theory can be formally
evaluated in what deserves to be called the Central Identity of Quantum Field Theory:
/integraldisplay
Dϕe−1
2ϕ.K.ϕ−V( ϕ) +J.ϕ=e−V( δ / δJ)e1
2J.K−1.J(1)
For any field theory we can always gather up all the fields, put them into one giant column
vector, and call the vector ϕ. We then single out the term quadratic in ϕwrite it as1
2ϕ.K.ϕ,
and call the rest V( ϕ) . I am using a compact notation in which spacetime coordinates and
any indices on the field, including Lorentz indices, are included in the indices of the formalmatrix K. We will often use (1) with V=0:
/integraldisplay
Dϕe−1
2ϕ.K.ϕ+J.ϕ=e1
2J.K−1.J(2)
But what if Kdoes not have an inverse?
This is not an esoteric phenomenon that occurs in some pathological field theory, but
in one of the most basic actions of physics, the Maxwell action
S(A)=/integraldisplay
d4xL=/integraldisplay
d4x/bracketleftBig
1
2Aμ(∂2gμν−∂μ∂ν)Aν+AμJμ/bracketrightBig
. (3)
The formal matrix Kin (2) is proportional to the differential operator (∂2gμν−∂μ∂ν)≡
Qμν. A matrix does not have an inverse if some of its eigenvalues are zero, that is, if when
acting on some vector, the matrix annihilates that vector. Well, observe that Qμνannihilates
vectors of the form ∂ν/Lambda1(x) :Qμν∂ν/Lambda1(x)=0. Thus Qμνhas no inverse.
There is absolutely nothing mysterious about this phenomenon; we have already en-
countered it in classical physics. Indeed, when we first learned the concepts of electricity,
we were told that only the “voltage drop” between two points has physical meaning. Ata more sophisticated level, we learned that we can always add any constant (or indeedany function of time) to the electrostatic potential (which is of course just “voltage”) since
III.4. Gauge Invariance | 183
by definition its gradient is the electric field. At an even more sophisticated level, we see
that solving Maxwell’s equation (which of course comes from just extremizing the action)amounts to finding the inverse Q
−1. [In the notation I am using here Maxwell’s equation
∂μFμν=Jνis written as QμνAν=Jμ, and the solution is Aν=(Q−1)νμJμ.]
Well,Q−1does not exist! What do we do? We learned that we must impose an additional
constraint on the gauge potential Aμ, known as “fixing a gauge.”
A mundane nonmystery
To emphasize the rather mundane nature of this gauge fixing problem (which some
older texts tend to make into something rather mysterious and almost hopelessly dif-ficult to understand), consider just an ordinary integral/integraltext
+∞
−∞dAe−A.K.A, with A=
(a,b)a 2-component vector and K=/parenleftBig10
00/parenrightBig
, a matrix without an inverse. Of course
you realize what the problem is: We have/integraltext+∞
−∞/integraltext+∞
−∞da db e−a2and the integral over
bdoes not exist. To define the integral we insert into it a delta function δ(b−ξ).
The integral becomes defined and actually does not depend on the arbitrary num-berξ. More generally, we can insert δ[f( b) ] with fsome function of our choice. In
the context of an ordinary integral, this procedure is of course ludicrous overkill, butwe will use the analog of this procedure in what follows. In this baby problem, wecould regard the variable b, and thus the integral over it, as “redundant.” As we will
see, gauge invariance is also a redundancy in our description of massless spin 1 parti-cles.
A massless spin 1 field is intrinsically different from a massive spin 1 field—that’s the
crux of the problem. The photon has only two polarization degrees of freedom. (You alreadylearned in classical electrodynamics that an electromagnetic wave has two transversedegrees of freedom.) This is the true physical origin of gauge invariance.
In this sense, gauge invariance is, strictly speaking, not a “real” symmetry but merely a
reflection of the fact that we used a redundant description: a Lorentz vector field to describetwo physical degrees of freedom.
Restricting the functional integral
I will now discuss the method for dealing with this redundancy invented by Faddeev and
Popov. As you will see presently, it is the analog of the method we used in our baby problem
above. Even in the context of electromagnetism this method is a bit of overkill, but it willprove to be essential for nonabelian gauge theories (as we will see in chapter VII.1) andfor gravity. I will describe the method using a completely general and somewhat abstract
language. In the next section, I will then apply the discussion here to a specific example.If you have some trouble with this section, you might find it helpful to go back and forthbetween the two sections.
184 | III. Renormalization and Gauge Invariance
Suppose we have to do the integral I≡/integraltext
DAeiS(A); this can be an ordinary integral
or a path integral. Suppose that under the transformation A→Agthe integrand and
the measure do not change, that is, S(A)=S(Ag)andDA=DAg. The transformations
obviously form a group, since if we transform again with g/prime, the integrand and the measure
do not change under the combined effect of gandg/primeandAg→(Ag)g/prime=Agg/prime. We would
like to write the integral Iin the form I=(/integraltext
Dg)J , with Jindependent of g. In other
words, we want to factor out the redundant integration over g. Note that Dg is the invariant
measure over the group of transformations and/integraltext
Dg is the volume of the group. Be aware
of the compactness of the notation in the case of a path integral: Aandgare both functions
of the spacetime coordinates x.
I want to emphasize that this hardly represents anything profound or mysterious. If
you have to do the integral I=/integraltext
dx dy eiS(x ,y)withS(x ,y)some function of x2+y2, you
know perfectly well to go to polar coordinates I=(/integraltext
dθ)J=(2π)J , where J=/integraltext
dr reiS(r)
is an integral over the radial coordinate ronly. The factor 2 πis precisely the volume of the
group of rotations in 2 dimensions.
Faddeev and Popov showed how to do this “going over to polar coordinates” in a
unified and elegant way. Following them, we first write the numeral “one” as1=/Delta1(A)/integraltext
Dgδ [f( A
g)], an equality that merely defines /Delta1(A) . Here fis some function
of our choice and /Delta1(A) , known as the Faddeev-Popov determinant, of course depends on
f. Next, note that [ /Delta1(Ag/prime)]−1=/integraltext
Dgδ [f( Ag/primeg)]=/integraltext
Dg/prime/primeδ[f( Ag/prime/prime)]=[/Delta1(A)]−1, where the
second equality follows upon defining g/prime/prime=g/primegand noting that Dg/prime/prime=Dg. In other words,
we showed that /Delta1(A) =/Delta1(A g): the Faddeev-Popov determinant is gauge invariant. We
now insert 1 into the integral Iwe have to do:
I=/integraldisplay
DAeiS(A)
=/integraldisplay
DAeiS(A)/Delta1(A)/integraldisplay
Dgδ [f( Ag)]
=/integraldisplay
Dg/integraldisplay
DAeiS(A)/Delta1(A)δ [f( Ag)] (4)
As physicists and not mathematicians, we have merrily interchanged the order of
integration.
At the physicist’s level of rigor, we are always allowed to change integration variables
until proven guilty. So let us change AtoAg−1; then
I=/parenleftbigg/integraldisplay
Dg/parenrightbigg/integraldisplay
DAeiS(A)/Delta1(A)δ [f (A)] (5)
where we have used the fact that DA ,S(A) , and/Delta1(A) are all invariant under A→Ag−1.
That’s it. We’ve done it. The group integration (/integraltext
Dg) has been factored out.
The volume of a compact group is finite, but in gauge theories there is a separate group
at every point in spacetime, and hence (/integraltext
Dg) is an infinite factor. (This also explains
why there is no gauge fixing problem in theories with global symmetries introduced in
III.4. Gauge Invariance | 185
chapter I.10.) Fortunately, in the path integral Zfor field theory we do not care about
overall factors in Z, as was explained in chapter I.3, and thus the factor (/integraltext
Dg) can simply
be thrown away.
Fixing the electromagnetic gauge
Let us now apply the Faddeev-Popov method to electromagnetism. The transformationleaving the action invariant is of course A
μ→Aμ−∂μ/Lambda1,s ogin the present context is
denoted by /Lambda1andAg≡Aμ−∂μ/Lambda1. Note also that since the integral Iwe started with is
independent of fit is still independent of fin spite of its appearance in (5). Choose
f (A)=∂A−σ, where σis a function of x. In particular, Iis independent of σand
so we can integrate Iwith an arbitrary functional of σ, in particular, the functional
e−(i/2ξ)/integraltext
d4xσ(x)2.
We now turn the crank. First, we calculate
[/Delta1(A)]−1≡/integraldisplay
Dgδ [f( Ag)]=/integraldisplay
D/Lambda1δ(∂A −∂2/Lambda1−σ) (6)
Next we note that in (5) /Delta1(A) appears multiplied by δ[f (A)] and so in evaluating [/Delta1(A)]−1
in (6) we can effectively set f (A)=∂A−σto zero. Thus from (6) we have /Delta1(A) “=”
[/integraltext
D/Lambda1δ(∂2/Lambda1)]−1. But this object does not even depend on A, so we can throw it away. Thus,
up to irrelevant overall factors that could be thrown away Iis just/integraltext
DAeiS(A)δ(∂A−σ).
Integrating over σ(x) as we said we were going to do, we finally obtain
Z=/integraldisplay
Dσe−(i/2ξ)/integraltext
d4xσ(x)2/integraldisplay
DAeiS(A)δ(∂A−σ)
=/integraldisplay
DAeiS(A)−(i/2 ξ)/integraltext
d4x(∂A)2(7)
Nifty trick by Faddeev and Popov, eh?
Thus, S(A) in (3) is effectively replaced by
Seff(A)=S(A)−1
2ξ/integraldisplay
d4x(∂A)2
=/integraldisplay
d4x/braceleftbigg1
2Aμ/bracketleftbigg
∂2gμν−/parenleftbigg
1−1
ξ/parenrightbigg
∂μ∂ν/bracketrightbigg
Aν+AμJμ/bracerightbigg
(8)
andQμνbyQμν
eff=∂2gμν−(1−1/ξ)∂μ∂νor in momentum space Qμν
eff=−k2gμν+(1−
1/ξ)kμkν, which does have an inverse. Indeed, you can check that
Qμν
eff/bracketleftbigg
−gνλ+(1−ξ)kνkλ
k2/bracketrightbigg1
k2=δμ
λ
Thus, the photon propagator can be chosen to be
(−i)
k2/bracketleftbigg
gνλ−(1−ξ)kνkλ
k2/bracketrightbigg
(9)
in agreement with the conclusion in chapter II.7.
186 | III. Renormalization and Gauge Invariance
While the Faddeev-Popov argument is a lot slicker, many physicists still prefer the explicit
Feynman argument given in chapter II.7. I do. When we deal with the Yang-Mills theoryand the Einstein theory, however, the Faddeev-Popov method is indispensable, as I havealready noted.
A photon can find no rest
Let us understand the physics behind the necessity for imposing by hand a (gauge fixing)constraint in gauge theories. In chapter I.5 we sidestepped this whole issue of fixing thegauge by treating the massive vector meson instead of the photon. In effect, we changedQ
μνto(∂2+m2)gμν−∂μ∂ν, which does have an inverse (in fact we even found the inverse
explicitly). We then showed that we could set the mass mto 0 in physical calculations.
There is, however, a huge and intrinsic difference between massive and massless
particles. Consider a massive particle moving along. We can always boost it to its restframe, or in more mathematical terms, we can always Lorentz transform the momentumof a massive particle to the reference momentum q
μ=m(1, 0, 0, 0 ). (As is the case
elsewhere in this book, if there is no risk of confusion, we write column vectors as rowvectors for typographical convenience.) To study the spin degrees of freedom, we shouldevidently sit in the rest frame of the particle and study how its states respond to rotation.The fancy pants way of saying this is we should study how the states of the particletransform under that particular subgroup of the Lorentz group (known as the little group)consisting of those Lorentz transformations /Lambda1that leave q
μinvariant, namely /Lambda1μ
νqν=qμ.
Forqμ=m(1, 0, 0, 0 ), the little group is obviously the rotation group SO( 3). We then apply
what we learned in nonrelativistic quantum mechanics and conclude that a spin jparticle
has(2j+1)spin states (or polarizations in classical physics), as already noted back in
chapter I.5.
But if the particle is massless, we can no longer find a Lorentz boost that would bring
us to its rest frame. A photon can find no rest!
For a massless particle, the best we can do is to transform the particle’s momentum to the
reference momentum qμ=ω(1, 0, 0, 1 )for some arbitrarily chosen ω. Again, this is just
a fancy way of saying that we can always call the direction of motion the third axis. What isthe little group that leaves q
μinvariant? Obviously, rotations around the third axis, forming
the group O(2), leave qμinvariant. The spin states of a massless particle of any spin around
its direction of motion are known as helicity states, as was already mentioned in chapter
II.1. For a particle of spin j, the helicities ±j are transformed into each other by parity
and time reversal, and thus both helicities must be present if the interactions the particleparticipates in respect these discrete symmetries, as is the case with the photon and thegraviton.
1In particular, the photon, as we have seen repeatedly, has only two polarization
degrees of freedom, instead of three, since we no longer have the full rotation group SO( 3).
1But not with the neutrino.
III.4. Gauge Invariance | 187
(You already learned in classical electrodynamics that an electromagnetic wave has two
transverse degrees of freedom.) For more on this, see appendix B.
In this sense, gauge invariance is strictly speaking not a “real” symmetry but merely a
reflection of the fact that we used a redundant description: we used a vector field Aμwith
its four degrees of freedom to describe two physical degrees of freedom. This is the truephysical origin of gauge invariance.
The condition /Lambda1
μ
νqν=qμshould leave us with a 3-parameter subgroup. To find the other
transformations, it suffices to look in the neighborhood of the identity, that is, at Lorentztransformations of the form /Lambda1(α ,β)=I+αA+βB+.... By inspection, we see that
A=⎛
⎜⎜⎜⎜⎜⎝010 0
100 −1
000 0
010 0⎞
⎟⎟⎟⎟⎟⎠=i(K
1+J2), B=⎛
⎜⎜⎜⎜⎜⎝001 0
000 0
100 −1
001 0⎞
⎟⎟⎟⎟⎟⎠=i(K
2−J1) (10)
where we used the notation for the generators of the Lorentz group from chapter II.3. Note
thatAandBare to a large extent determined by the fact that JandKare symmetric and
antisymmetric, respectively.
By direct computation or by invoking the celebrated minus sign in (II.3.9), we find
that [A,B]=0. Also, [J3,A]=Band [J3,B]=−Aso that, as expected, (A,B)form a
2-component vector under O(2)rotations around the third axis. (For those who must
know, the generators A,B, andJ3generate the group ISO( 2), the invariance group of
the Euclidean 2-plane, consisting of two translations and one rotation.)
The preceding paragraph establishing the little group for massless particles applies
for any spin, including zero. Now specialize to a spin 1 massless particle with the twopolarization vectors /epsilon1
±(q)=(1/√
2)(0, 1, ±i,0). The polarization vectors are defined by
how they transform under rotation eiφJ 3. So it is natural to ask how /epsilon1±(q)transform under
/Lambda1(α ,β). Inspecting (10), we see that
/epsilon1±(q)→/epsilon1±(q)+1√
2(α±iβ)q (11)
We recognize (11) as a gauge transformation (as was explained in chapter II.7). For a mass-
less spin 1 particle, the gauge transformation is contained in the Lorentz transformations!
Suppose we construct the corresponding spin 1 field as in chapter II.5 [and in analogy
to (I.8.11) and (II.2.10)]:
Aμ(x)=/integraldisplayd3k/radicalbig
(2π)32ωk/summationdisplay
α=1, 2[a(α)(/vectork)ε(α)
μ(k)e−i(ω kt−/vectork./vectorx)+a†(α)(/vectork)ε∗(α)
μ(k)ei(ωkt−/vectork./vectorx))] (12)
withωk=|/vectork|. The polarization vectors ε(α)
μ(k)are of coursed determined by the condition
kμε(α)
μ(k)=0, which we could easily satisfy by defining ε(α)(k)=/Lambda1(q→k)ε(α)(q), where
/Lambda1(q→k)denotes a Lorentz transformation that brings the reference momentum qtok.
Note that the ε(α)(k) thus constructed has a vanishing time component. (To see this,
first boost ε(α)
μ(q) along the third axis and then rotate, for example.) Hence, kμε(α)
μ(k)=
−/vectork./vectorε(α)(k)=0. These properties of ε(α)
μ(k) translate into A0(x)=0 and− →∇./vectorA(x)=0.
188 | III. Renormalization and Gauge Invariance
These two constraints cut the four degrees of freedom contained in Aμ(x) down to two
and fix what is known as the Coulomb or radiation gauge.2
Given the enormous importance of gauge invariance, it might be instructive to review
the logic underlying the “poor man’s approach” to gauge invariance (which, as I mentionedin chapter I.5, I learned from Coleman) adopted in this book for pedagogical reasons. Youcould have fun faking Feynman’s mannerism and accent, saying, “Aw shucks, all that fancytalk about little groups! Who needs it? Those experimentalists won’t ever be able to provethat the photon mass is mathematically zero anyway.”
So start, as in chapter I.5, with the two equations needed for describing a spin 1 massive
particle:
(∂2+m2)Aμ=0 (13)
and
∂μAμ=0 (14)
Equation (14) is needed to cut the number of degrees of freedom contained in Aμdown
from four to three.
Lo and behold, (13) and (14) are equivalent to the single equation
∂μ(∂μAν−∂νAμ)+m2Aν=0 (15)
Obviously, (13) and (14) together imply (15). To verify that (15) implies (13) and (14), we
act with ∂νon (15) and obtain
m2∂A=0 (16)
which for m/negationslash=0 requires ∂A=0, namely (14). Plugging this into (15) we obtain (13).
Having packaged two equations into one, we note that we can derive this single equation
(15) by varying the Lagrangian
L=−1
4FμνFμν+1
2m2A2(17)
withFμν≡∂μAν−∂νAμ.
Next, suppose we include a source Jμfor this particle by changing the Lagrangian to
L=−1
4FμνFμν+1
2m2A2+AμJμ(18)
with the resulting equation of motion
∂μ(∂μAν−∂νAμ)+m2Aν=−Jν (19)
But now observe that when we act with ∂νon (19) we obtain
m2∂A=−∂J (20)
2For a much more detailed and leisurely discussion, see S. Weinberg, Quantum Theory of Fields, pp. 69–74
and 246–255.
III.4. Gauge Invariance | 189
We recover (14) only if ∂μJμ=0, that is, if the source producing the particle, commonly
know as the current, is conserved.
Put more vividly, suppose the experimentalists who constructed the accelerator (or
whatever) to produce the spin 1 particle messed up and failed to insure that ∂μJμ=0;
then∂A/negationslash=0 and a spin 0 excitation would also be produced. To make sure that the beam
of spin 1 particles is not contaminated with spin 0 particles, the accelerator builders mustassure us that the source J
μin the Lagrangian (18) is indeed conserved.
Now, if we want to study massless spin 1 particles, we simply set m=0 in (18). The
“poor man” ends up (just like the “rich man”) using the Lagrangian
L=−1
4FμνFμν+AμJμ(21)
to describe the photon. Lo and behold (as we exclaimed in chapter II.7), Lis left invariant
by the gauge transformation Aμ→Aμ−∂μ/Lambda1for any /Lambda1(x). (As was also explained in that
chapter, the third polarization decouples in the limit m→0.) The “poor man” has thus
discovered gauge invariance!
However, as I warned in chapter I.5, depending on his or her personality, the poor man
could also wake up in the middle of the night worrying that physics might be discontinuousin the limit m→0. Thus the little group discussion is needed to remove that nightmare.
But then a “real” physicist in the Feynman mode could always counter that for any physicalmeasurement everything must be okay as long as the duration of the experiment is shortcompared to the characteristic time 1 /m. More on this issue in chapter VIII.1.
A reflection on gauge symmetry
As we will see later and as you might have heard, much of the world beyond electro-
magnetism is also described by gauge theories. But as we saw here, gauge theories arealso deeply disturbing and unsatisfying in some sense: They are built on a redundancyof description. The electromagnetic gauge transformation A
μ→Aμ−∂μ/Lambda1is not truly a
symmetry stating that two physical states have the same properties. Rather, it tells us thatthe two gauge potentials A
μandAμ−∂μ/Lambda1describe the same physical state. In your or-
derly study of physics, the first place where Aμbecomes indispensable is the Schr ¨odinger
equation, as I will explain in chapter IV .4. Within classical physics, you got along perfectlywell with just /vectorEand/vectorB. Some physicists are looking for a formulation of quantum elec-
trodynamics without using A
μ, but so far have failed to turn up an attractive alternative
to what we have. It is conceivable that a truly deep advance in theoretical physics wouldinvolve writing down quantum electrodynamics without writing A
μ.
III.5 Field Theory without Relativity
Slower in its maturity
Quantum field theory at its birth was relativistic. Later in its maturity, it found applications
in condensed matter physics. We will have a lot more to say about the role of quantum fieldtheory in condensed matter, but for now, we have the more modest goal of learning howto take the nonrelativistic limit of a quantum field theory.
The Lorentz invariant scalar field theory
L=(∂/Phi1†)(∂/Phi1) −m2/Phi1†/Phi1−λ(/Phi1†/Phi1)2(1)
(withλ>0 as always) describes a bunch of interacting bosons. It should certainly contain
the physics of slowly moving bosons. For clarity consider first the relativistic Klein-Gordonequation
(∂2+m2)/Phi1=0 (2)
for a free scalar field. A mode with energy E=m+εwould oscillate in time as /Phi1∝e−iEt.
In the nonrelativistic limit, the kinetic energy εis much smaller than the rest mass
m. It makes sense to write /Phi1(/vectorx,t)=e−imtϕ(/vectorx,t), with the field ϕoscillating in time
much more slowly than e−imt. Plugging into (2) and using the identity (∂/∂t)e−imt(...)=
e−imt(−im +∂/∂t)( ...)twice, we obtain (−im +∂/∂t)2ϕ−/vector∇2ϕ+m2ϕ=0. Dropping the
term(∂2/∂t2)ϕas small compared to −2im(∂/∂t)ϕ , we find Schr ¨odinger’s equation, as we
had better:
i∂
∂tϕ=−/vector∇2
2mϕ (3)
By the way, the Klein-Gordon equation was actually discovered before Schr ¨odinger’s
equation.
III.5. Nonrelativistic Field Theory | 191
Having absorbed this, you can now easily take the nonrelativistic limit of a quantum
field theory. Simply plug
/Phi1(/vectorx,t)=1√
2me−imtϕ(/vectorx,t) (4)
into (1) . (The factor 1 /√
2mis for later convenience.) For example,
∂/Phi1†
∂t∂/Phi1
∂t−m2/Phi1†/Phi1→1
2m/braceleftbigg/bracketleftbigg/parenleftbigg
im+∂
∂t/parenrightbigg
ϕ†/bracketrightbigg/bracketleftbigg /parenleftbigg
−im+∂
∂t/parenrightbigg
ϕ/bracketrightbigg
−m2ϕ†ϕ/bracerightbigg
/similarequal1
2i/parenleftBigg
ϕ†∂ϕ
∂t−∂ϕ†
∂tϕ/parenrightBigg
(5)
After an integration by parts we arrive at
L=iϕ†∂0ϕ−1
2m∂iϕ†∂iϕ−g2(ϕ†ϕ)2(6)
where g2=λ/4m2.
As we saw in chapter I.10 the theory (1) enjoys a conserved Noether current Jμ=
i(/Phi1†∂μ/Phi1−∂μ/Phi1†/Phi1). The density J0reduces to ϕ†ϕ, precisely as you would expect, while
Jireduces to (i/2m)(ϕ†∂iϕ−∂iϕ†ϕ). When you first took a course in quantum mechanics,
didn’t you wonder why the density ρ≡ϕ†ϕand the current Ji=(i/2m)(ϕ†∂iϕ−∂iϕ†ϕ)
look so different? As to be expected, various expressions inevitably become uglier whenreduced from a more symmetric to a less symmetric theory.
Number is conjugate to phase angle
Let me point out some differences between the relativistic and nonrelativistic case.
The most striking is that the relativistic theory is quadratic in time derivative, while
the nonrelativistic theory is linear in time derivative. Thus, in the nonrelativistic theorythe momentum density conjugate to the field ϕ, namely δL/δ∂
0ϕ, is just iϕ†, so that
[ϕ†(/vectorx,t),ϕ(/vectorx/prime,t)]=−δ(D)(/vectorx−/vectorx/prime). In condensed matter physics it is often illuminating to
writeϕ=√ρeiθso that
L=i
2∂0ρ−ρ∂0θ−1
2m/bracketleftbigg
ρ(∂iθ)2+1
4ρ(∂iρ)2/bracketrightbigg
−g2ρ2(7)
The first term is a total divergence. The second term tells us something of great impor-
tance1in condensed matter physics: in the canonical formalism (chapter I.8), the momen-
tum density conjugate to the phase field θ(x) isδL/δ∂ 0θ=−ρand thus Heisenberg tells
us that
[ρ(/vectorx,t),θ(/vectorx/prime,t)]=iδ(D)(/vectorx−/vectorx/prime) (8)
1See P. Anderson, Basic Notions of Condensed Matter Physics , p. 235.
192 | III. Renormalization and Gauge Invariance
Integrating and defining N≡/integraltext
dDxρ(/vectorx,t)=the total number of bosons, we find one of
the most important relations in condensed matter physics
[N,θ]=i (9)
Number is conjugate to phase angle, just as momentum is conjugate to position. Marvel
at the elegance of this! You would learn in a condensed matter course that this fundamentalrelation underlies the physics of the Josephson junction.
You may know that a system of bosons with a “hard core” repulsion between them is a
superfluid at zero temperature. In particular, Bogoliubov showed that the system containsan elementary excitation obeying a linear dispersion relation.
2I will discuss superfluidity
in chapter V .1.
In the path integral formalism, going from the complex field ϕ=ϕ1+iϕ2toρand
θamounts to a change of integration variables, as I remarked back in chapter I.8. In the
canonical formalism, since one deals with operators, one has to tread with somewhat morefinesse.
The sign of repulsion
In the nonrelativistic theory (7) it is clear that the bosons repel each other: Piling particlesinto a high density region would cost you an energy density g
2ρ2. But it is less clear in
the relativistic theory that λ(/Phi1†/Phi1)2withλpositive corresponds to repulsion. I outline one
method in exercise III.5.3, but here let’s just take a flying heuristic guess. The Hamiltonian(density) involves the negative of the Lagrangian and hence goes as λ(/Phi1
†/Phi1)2for large /Phi1
and would thus be unbounded below for λ<0. We know physically that a free Bose gas
tends to condense and clump, and with an attractive interaction it surely might want tocollapse. We naturally guess that λ>0 corresponds to repulsion.
I next give you a more foolproof method. Using the central identity of quantum field
theory we can rewrite the path integral for the theory in (1) as
Z=/integraldisplay
D/Phi1Dσei/integraltext
d4x[(∂/Phi1†)(∂/Phi1)−m2/Phi1†/Phi1+2σ/Phi1†/Phi1+(1/λ)σ2](10)
Condensed matter physicists call the transformation from (1) to the Lagrangian L=
(∂/Phi1†)(∂/Phi1) −m2/Phi1†/Phi1+2σ/Phi1†/Phi1+(1/λ)σ2the Hubbard-Stratonovich transformation. In
field theory, a field that does not have kinetic energy, such as σ, is known as an auxiliary field
and can be integrated out in the path integral. When we come to the superfield formalismin chapter VIII.4, auxiliary fields will play an important role.
Indeed, you might recall from chapter III.2 how a theory with an intermediate vector
boson could generate Fermi’s theory of the weak interaction. The same physics is involvedhere: The theory (10) in which the /Phi1field is coupled to an “intermediate σboson” can
generate the theory (1).
2For example, L.D. Landau and E. M. Lifschitz, Statisical Physics , p. 238.
III.5. Nonrelativistic Field Theory | 193
Ifσwere a “normal scalar field” of the type we have studied, that is, if the terms quadratic
inσin the Lagrangian had the form1
2(∂σ)2−1
2M2σ2, then its propagator would be
i/(k2−M2+iε). The scattering amplitude between two /Phi1bosons would be proportional
to this propagator. We learned in chapter I.4 that the exchange of a scalar field leads to anattractive force.
Butσis not a normal field as evidenced by the fact that the Lagrangian contains
only the quadratic term +(1/λ)σ
2. Thus its propagator is simply i/(1/λ)=iλ, which (for
λ>0)has a sign opposite to the normal propagator evaluated at low-momentum transfer
i/(k2−M2+iε)/similarequal−i/M2. We conclude that σexchange leads to a repulsive force.
Incidentally, this argument also shows that the repulsion is infinitely short ranged, like
a delta function interaction. Normally, as we learned in chapter I.4 the range is determinedby the interplay between the k
2and the M2terms. Here the situation is as if the M2term
is infinitely large. We can also argue that the interaction λ(/Phi1†/Phi1)2involves creating two
bosons and then annihilating them both at the same spacetime point.
Finite density
One final point of physics that people trained as particle physicists do not always remem-ber: Condensed matter physicists are not interested in empty space, but want to have afinite density ¯ρof bosons around. We learned in statistical mechanics to add a chemical
potential term μϕ
†ϕto the Lagrangian (6). Up to an irrelevant (in this context!) additive
constant, we can rewrite the resulting Lagrangian as
L=iϕ†∂0ϕ−1
2m∂iϕ†∂iϕ−g2(ϕ†ϕ−¯ρ)2(11)
Amusingly, mass appears in different places in relativistic and nonrelativistic field
theories. To proceed further, I have to develop the concept of spontaneous symmetrybreaking. Thus, adios for now. We will come back to superfluidity in due time.
Exercises
III.5.1 Obtain the Klein-Gordon equation for a particle in an electrostatic potential (such as that of the nucleus)
by the gauge principle of replacing (∂/∂t) in (2) by ∂/∂t−ieA 0. Show that in the nonrelativistic limit
this reduces to the Schr ¨odinger’s equation for a particle in an external potential.
III.5.2 Take the nonrelativistic limit of the Dirac Lagrangian.
III.5.3 Given a field theory we can compute the scattering amplitude of two particles in the nonrelativistic limit.
We then postulate an interaction potential U(/vectorx)between the two particles and use nonrelativistic quan-
tum mechanics to calculate the scattering amplitude, for example in Born approximation. Comparingthe two scattering amplitudes we can determine U(/vectorx). Derive the Yukawa and the Coulomb potentials
this way. The application of this method to the λ(/Phi1
†/Phi1)2interaction is slightly problematic since the
delta function interaction is a bit singular, but it should be all right for determining whether the force isrepulsive or attractive.
III.6 The Magnetic Moment of the Electron
Dirac’s triumph
I said in the preface that the emphasis in this book is not on computation, but how can I
not tell you about the greatest triumph of quantum field theory?
After Dirac wrote down his equation, the next step was to study how the electron interacts
with the electromagnetic field. According to the gauge principle already used to writethe Schr ¨odinger’s equation in an electromagnetic field, to obtain the Dirac equation for
an electron in an external electromagnetic field we merely have to replace the ordinaryderivative ∂
μby the covariant derivative Dμ=∂μ−ieAμ:
(iγμDμ−m)ψ=0 (1)
Recall (II.1.27).
Acting on this equation with (iγμDμ+m), we obtain −(γμγνDμDν+m2)ψ=0. We
haveγμγνDμDν=1
2({γμ,γν}+[γμ,γν])DμDν=DμDμ−iσμνDμDνandiσμνDμDν=
(i/2)σμν[Dμ,Dν]=(e/2)σμνFμν. Thus
/parenleftbigg
DμDμ−e
2σμνFμν+m2/parenrightbigg
ψ=0 (2)
Now consider a weak constant magnetic field pointing in the 3rd direction for definite-
ness, weak so that we can ignore the ( Ai)2term in ( Di)2. By gauge invariance, we can
choose A0=0,A1=−1
2Bx2, andA2=1
2Bx1(so that F12=∂1A2−∂2A1=B). As we will
see, this is one calculation in which we really have to keep track of factors of 2. Then
(Di)2=(∂i)2−ie(∂iAi+Ai∂i)+O(A2
i)
=(∂i)2−2ie
2B(x1∂2−x2∂1)+O(A2
i)
=/vector∇2−e/vectorB./vectorx×/vectorp+O(A2
i) (3)
Note that we used ∂iAi+Ai∂i=(∂iAi)+2Ai∂i=2Ai∂i, where in (∂iAi)the partial deriva-
tive acts only on Ai. You may have recognized /vectorL≡/vectorx×/vectorpas the orbital angular momentum
III.6. Magnetic Moment of Electron | 195
operator. Thus, the orbital angular momentum generates an orbital magnetic moment that
interacts with the magnetic field.
This calculation makes good physical sense. If we were studying the interaction of a
charged scalar field /Phi1with an external electromagnetic field we would start with
(DμDμ+m2)/Phi1=0 (4)
obtained by replacing the ordinary derivative in the Klein-Gordon equation by covariant
derivatives. We would then go through the same calculation as in (3). Comparing (4) with(2) we see that the spin of the electron contributes the additional term (e/2)σ
μνFμν.
As in chapter II.1 we write ψ=/parenleftBigφ
χ/parenrightBig
in the Dirac basis and focus on φsince in the
nonrelativistic limit it dominates χ. Recall that in that basis σij=εijk/parenleftBigσk0
0σk/parenrightBig
. Thus
(e/2)σμνFμνacting on φis effectively equal to (e/2)σ3(F12−F21)=(e/2)2σ3B=2e/vectorB./vectorS
since /vectorS=(/vectorσ/2). Make sure you understand all the factors of 2! Meanwhile, according to
what I told you in chapter II.1, we should write φ=e−imt/Psi1, where /Psi1oscillates much more
slowly than e−imtso that (∂2
0+m2)e−imt/Psi1/similarequale−imt[−2im(∂/∂t)/Psi1 ]. Putting it all together,
we have
/bracketleftbigg
−2im∂
∂t−/vector∇2−e/vectorB.(/vectorL+2/vectorS)/bracketrightbigg
/Psi1=0 (5)
There you have it! As if by magic, Dirac’s equation tells us that a unit of spin angular
momentum interacts with a magnetic field twice as much as a unit of orbital angularmomentum, an observational fact that had puzzled physicists deeply at the time. Thecalculation leading to (5) is justly celebrated as one of the greatest in the history of physics.
The story is that Dirac did not do this calculation until a day after he discovered his
equation, so sure was he that the equation had to be right. Another version is that hedreaded the possibility that the magnetic moment would come out wrong and that Naturewould not take advantage of his beautiful equation.
Another way of seeing that the Dirac equation contains a magnetic moment is by the
Gordon decomposition, the proof of which is given in an exercise:
¯u(p/prime)γμu(p)=¯u(p/prime)/bracketleftbigg(p/prime+p)μ
2m+iσμν(p/prime−p)ν
2m/bracketrightbigg
u(p) (6)
Looking at the interaction with an electromagnetic field ¯u(p/prime)γμu(p)Aμ(p/prime−p), we see
that the first term in (6) only depends on the momentum (p/prime+p)μand would have
been there even if we were treating the interaction of a charged scalar particle with the
electromagnetic field to first order. The second term involves spin and gives the mag-netic moment. One way of saying this is that ¯u(p
/prime)γμu(p) contains a magnetic moment
component.
196 | III. Renormalization and Gauge Invariance
The anomalous magnetic moment
With improvements in experimental techniques, it became clear by the late 1940’s that the
magnetic moment of the electron was larger than the value calculated by Dirac by a factorof 1.00118 ±0.00003. The challenge to any theory of quantum electrodynamics was to
calculate this so-called anomalous magnetic moment. As you probably know, Schwinger’sspectacular success in meeting this challenge established the correctness of relativisticquantum field theory, at least in dealing with electromagnetic phenomena, beyond anydoubt.
Before we plunge into the calculation, note that Lorentz invariance and current conser-
vation tell us (see exercise III.6.3) that the matrix element of the electromagnetic currentmust have the form (here |p,s/angbracketrightdenotes a state with an electron of momentum pand
polarization s)
/angbracketleftp/prime,s/prime|Jμ(0)|p,s/angbracketright=¯u(p/prime,s/prime)/bracketleftbigg
γμF1(q2)+iσμνqν
2mF2(q2)/bracketrightbigg
u(p ,s) (7)
where q≡(p/prime−p). The functions F1(q2)andF2(q2), about which Lorentz invariance can
tell us nothing, are known as form factors. To leading order in momentum transfer q, (7)
becomes
¯u(p/prime,s/prime)/braceleftbigg(p/prime+p)μ
2mF1(0)+iσμνqν
2m[F1(0)+F2(0)]/bracerightbigg
u(p ,s)
by the Gordon decomposition. The coefficient of the first term is the electric charge
observed by experimentalists and is by definition equal to 1. (To see this, think of potentialscattering, for example. See chapter II.6.) Thus F
1(0)=1. The magnetic moment of the
electron is shifted from the Dirac value by a factor 1 +F2(0).
Schwinger’s triumph
Let us now calculate F2(0)to order α=e2/4π. First draw all the relevant Feynman dia-
grams to this order (fig. III.6.1). Except for figure 1b, all the Feynman diagrams are clearlyproportional to ¯u(p
/prime,s/prime)γμu(p ,s)and thus contribute to F1(q2), which we don’t care about.
Happy are we! We only have to calculate one Feynman diagram.
pqp’
pqp’
p’ + k
p + kk
(a) (b) (c) (d) (e)
Figure III.6.1
III.6. Magnetic Moment of Electron | 197
It is convenient to normalize the contribution of figure 1b by comparing it to the lowest
order contribution of figure 1a and write the sum of the two contributions as ¯u(γμ+/Gamma1μ)u.
Applying the Feynman rules, we find
/Gamma1μ=/integraldisplayd4k
(2π)4−i
k2/parenleftbigg
ieγν i
/negationslashp/prime+ /negationslashk−mγμ i
/negationslashp+ /negationslashk−mieγν/parenrightbigg
(8)
I will now go through the calculation in some detail not only because it is important, but
also because we will be using a variety of neat tricks springing from the brilliant minds ofSchwinger and Feynman. You should verify all the steps of course.
Simplifying somewhat we obtain /Gamma1
μ=−ie2/integraltext
[d4k/(2π)4](Nμ/D), where
Nμ=γν(/negationslashp/prime+ /negationslashk+m)γμ(/negationslashp+ /negationslashk+m)γν (9)
and
1
D=1
(p/prime+k)2−m21
(p+k)2−m21
k2=2/integraldisplay
dα dβ1
D. (10)
We have used the identity (D.16). The integral is evaluated over the triangle in the ( α-β)
plane bounded by α=0,β=0, and α+β=1, and
D=[k2+2k(αp/prime+βp) ]3=[l2−(α+β)2m2]3+O(q2) (11)
where we completed a square by defining k=l−(αp/prime+βp) . The momentum integration
is now over d4l.
Our strategy is to massage Nμinto a form consisting of a linear combination of γμ,pμ,
andp/primeμ. Invoking the Gordon decomposition (6) we can write (7) as
¯u/braceleftbigg
γμ[F1(q2)+F2(q2)]−1
2m(p/prime+p)μF2(q2)/bracerightbigg
u
Thus, to extract F2(0)we can throw away without ceremony any term proportional to γμ
that we encounter while massaging Nμ. So, let’s proceed.
Eliminating kin favor of lin (9) we obtain
Nμ=γν[/negationslashl+ /negationslashP/prime+m]γμ[/negationslashl+ /negationslashP+m]γν (12)
where P/primeμ≡(1−α)p/primeμ−βpμandPμ≡(1−β)pμ−αp/primeμ. I will use the identities in
appendix D repeatedly, without alerting you every time I use one. It is convenient toorganize the terms in N
μby powers of m. (Here I give up writing in complete grammatical
sentences.)
1. The m2term: a γμterm, throw away.
2. The mterms: organize by powers of l. The term linear in lintegrates to 0 by symmetry.
Thus, we are left with the term independent of l:
m(γν/negationslashP/primeγμγν+γνγμ/negationslashPγν)=4m[(1−2α)p/primeμ+(1−2β)pμ]
→4m(1−α−β)(p/prime+p)μ(13)
In the last step I used a handy trick; since Dis symmetric under α←→β, we can sym-
metrize the terms we get in Nμ.
198 | III. Renormalization and Gauge Invariance
3. Finally, the most complicated m0term. The term quadratic in l: note that we can effectively
replace lσlτinside/integraltext
d4l/(2π)4by1
4ηστl2by Lorentz invariance (this step is possible because
we have shifted the integration variable so that Dis a Lorentz invariant function of l2.) Thus,
the term quadratic in lgives rise to a γμterm. Throw it away. Again we throw away the term
linear in l, leaving [use (D.6) here!]
γν/negationslashP/primeγμ/negationslashPγν=− 2/negationslashPγμ/negationslashP/prime
→− 2[(1−β)/negationslashp−αm]γμ[(1−α)/negationslashp/prime−βm] (14)
where in the last step we remembered that /Gamma1μis to be sandwiched between ¯u(p/prime)andu(p) .
Again, it is convenient to organize the terms in (14) by powers of m. With the various tricks
we have already used, we find that the m2term can be thrown away, the mterm gives
2m(p/prime+p)μ[α(1−α)+β(1−β)], and the m0term gives 2 m(p/prime+p)μ[−2(1−α)(1−β)].
Putting it altogether, we find that Nμ→2m(p/prime+p)μ(α+β)(1−α−β)
We can now do the integral/integraltext
[d4l/(2π)4](1/D)using (D.11). Finally, we obtain
/Gamma1μ=− 2ie2/integraldisplay
dα dβ(−i
32π2)1
(α+β)2m2Nμ
=−e2
8π21
2m(p/prime+p)μ(15)
and thus, trumpets please:
F2(0)=e2
8π2=α
2π(16)
Schwinger’s announcement of this result in 1948 had an electrifying impact on the theo-
retical physics community.
I gave you in this chapter not one, but two, of the great triumphs of twentieth century
physics, although admittedly the first is not a result of field theory per se.
Exercises
III.6.1 Evaluate ¯u(p/prime)(/negationslashp/primeγμ+γμ/negationslashp)u(p) in two different ways and thus prove Gordon decomposition.
III.6.2 Check that (7) is consistent with current conservation. [Hint: By translation invariance (we suppress the
spin variable)
/angbracketleftp/prime|Jμ(x)|p/angbracketright=/angbracketleftp/prime|Jμ(0)|p/angbracketrightei(p/prime−p)x
and hence
/angbracketleftp/prime|∂μJμ(x)|p/angbracketright=i(p/prime−p)μ/angbracketleftp/prime|Jμ(0)|p/angbracketrightei(p/prime−p)x
Thus current conservation implies that qμ/angbracketleftp/prime|Jμ(0)|p/angbracketright=0.]
III.6.3 By Lorentz invariance the right hand side of (7) has to be a vector. The only possibilities are ¯uγμu,
(p+p/prime)μ¯uu, and(p−p/prime)μ¯uu. The last term is ruled out because it would not be consistent with current
conservation. Show that the form given in (7) is in fact the most general allowed.
III.6. Magnetic Moment of Electron | 199
III.6.4 In chapter II.6, when discussing electron-proton scattering, we ignored the strong interaction that the
proton participates in. Argue that the effects of the strong interaction could be included phenomenolog-ically by replacing the vertex ¯u(P ,S)γ
μu(p ,s)in (II.6.1) by
/angbracketleftP,S|Jμ(0)|p,s/angbracketright=¯u(P ,S)/bracketleftbigg
γμF1(q2)+iσμνqν
2mF2(q2)/bracketrightbigg
u(p ,s) (17)
Careful measurements of electron-proton scattering, thus determining the two proton form factors
F1(q2)andF2(q2), earned R. Hofstadter the 1961 Nobel Prize. While we could account for the general
behavior of these two form factors, we are still unable to calculate them from first principles (in contrastto the corresponding form factors for the electron.) See chapters IV .2 and VII.3.
III.7Polarizing the Vacuum and
Renormalizing the Charge
A photon can fluctuate into an electron and a positron
One early triumph of quantum electrodynamics is the understanding of how quantum
fluctuations affect the way the photon propagates. A photon can always metamorphose intoan electron and a positron that, after a short time mandated by the uncertainty principle,annihilate each other becoming a photon again. The process, which keeps on repeatingitself, is depicted in figure III.7.1.
Quantum fluctuations are not limited to what we just described. The electron and
positron can interact by exchanging a photon, which in turn can change into an electronand a positron, and so on and so forth. The full process is shown in figure III.7.2, where theshaded object, denoted by i/Pi1
μν(q) and known as the vacuum polarization tensor, is given
by an infinite number of Feynman diagrams, as shown in figure III.7.3. Figure III.7.1 isobtained from figure III.7.2 by approximating i/Pi1
μν(q) by its lowest order diagram.
It is convenient to rewrite the Lagrangian L=¯ψ[iγμ(∂μ−ieAμ)−m]ψ−1
4FμνFμνby
letting A→(1/e)A, which we are always allowed to do, so that
L=¯ψ[iγμ(∂μ−iAμ)−m]ψ−1
4e2FμνFμν(1)
Note that the gauge transformation leaving Linvariant is given by ψ→eiαψandAμ→
Aμ+∂μα. The photon propagator (chapter III.4), obtained roughly speaking by inverting
(1/4e2)FμνFμν, is now proportional to e2:
iDμν(q)=−ie2
q2/bracketleftbigg
gμν−(1−ξ)qμqν
q2/bracketrightbigg
(2)
Every time a photon is exchanged, the amplitude gets a factor of e2. This is just a trivial
but convenient change and does not affect the physics in the slightest. For example, in the
Feynman diagram we calculated in chapter II.6 for electron-electron scattering, the factore
2can be thought of as being associated with the photon propagator rather than as coming
from the interaction vertices. In this interpretation e2measures the ease with which the
III.7. Polarizing the Vacuum | 201
+ +
+
Figure III.7.1
+ +
+
Figure III.7.2
=
+ ++
++
pp + q
Figure III.7.3
photon propagates through spacetime. The smaller e2, the more action it takes to have
the photon propagate, and the harder for the photon to propagate, the weaker the effect ofelectromagnetism.
The diagrammatic proof of gauge invariance given in chapter II.7 implies that
q
μ/Pi1μν(q)=0. T ogether with Lorentz invariance, this requires that
/Pi1μν(q)=(qμqν−gμνq2)/Pi1(q2) (3)
The physical or renormalized photon propagator as shown in figure III.7.2 is then given
by the geometric series
iDP
μν(q)=iDμν(q)+iDμλ(q)i/Pi1λρ(q)iD ρν(q)
+iDμλ(q)i/Pi1λρ(q)iDρσ(q)i/Pi1σκ(q)iDκν(q)+...
=−ie2
q2gμν{1−e2/Pi1(q2)+[e2/Pi1(q2)]2+...}+q μqνterm
=−ie2
q2gμν1
1+e2/Pi1(q2)+qμqνterm (4)
Because of (3) the (1−ξ)(qμqλ/q2)part of Dμλ(q) is annihilated when it encounters
/Pi1λρ(q). Thus, in iDP
μν(q) the gauge parameter ξenters only into the qμqνterm and drops
out in physical amplitudes, as explained in chapter II.7.
202 | III. Renormalization and Gauge Invariance
The residue of the pole in iDP
μν(q) is the physical or renormalized charge squared:
e2
R=e2 1
1+e2/Pi1(0)(5)
Respect for gauge invariance
In order to determine eRin terms of e, let us calculate to lowest order
i/Pi1μν(q)=(−)/integraldisplayd4p
(2π)4tr/parenleftbigg
iγν i
/negationslashp+ /negationslashq−miγμi
/negationslashp−m/parenrightbigg
(6)
For large pthe integrand goes as 1 /p2with a subleading term going as m2/p4causing
the integral to have a quadratically divergent and a logarithmically divergent piece. (Yousee, it is easy to slip into bad language.) Not a conceptual problem at all, as I explainedin chapter III.1. We simply regularize. But now there is a delicate point: Since gaugeinvariance plays a crucial role, we must make sure that our regularization respects gaugeinvariance.
In the Pauli-Villars regularization (III.1.13) we replace (6) by
i/Pi1μν(q)=(−)/integraldisplayd4p
(2π)4/bracketleftbigg
tr/parenleftbigg
iγν i
/negationslashp+ /negationslashq−miγμi
/negationslashp−m/parenrightbigg
−/summationdisplay
acatr/parenleftbigg
iγν i
/negationslashp+ /negationslashq−maiγμi
/negationslashp−ma/parenrightbigg/bracketrightBigg
(7)
Now the integrand goes as (1−/summationtext
aca)(1/p2)with a subleading term going as (m2−/summationtext
acam2
a)(1/p4), and thus the integral would converge if we choose caandmasuch that
/summationdisplay
aca=1 (8)
and
/summationdisplay
acam2
a=m2(9)
Clearly, we have to introduce at least two regulator masses. We are confessing to ignorance
of the physics above the mass scale ma. The integral in (7) is effectively cut off when the
momentum pexceeds ma.
Does a bell ring for you? It should, as this discussion conceptually parallels that in the
appendix to chapter I.9.
The gauge invariant form (3) we expect to get actually suggests that we need fewer
regulator terms than we think. Imagine expanding (6) in powers of q. Since
/Pi1μν(q)=(qμqν−gμνq2)[/Pi1(0)+...]
we are only interested in terms of O(q2)and higher in the Feynman integral. If we expand
the integrand in (6), we see that the term of O(q2)goes as 1 /p4for large p, thus giving a
logarithmically divergent (speaking bad language again!) contribution. (Incidentally, you
III.7. Polarizing the Vacuum | 203
may recall that this sort of argument was also used in chapter III.3.) It seems that we need
only one regulator. This argument is not rigorous because we have not proved that /Pi1(q2)
has a power series expansion in q2, but instead of worrying about it let us proceed with
the calculation.
Once the integral is convergent, the proof of gauge invariance given in chapter II.7 now
goes through. Let us recall briefly how the proof went. In computing qμ/Pi1μν(q) we use the
identity
1
/negationslashp+ /negationslashq−m/negationslashq1
/negationslashp−m=1
/negationslashp−m−1
/negationslashp+ /negationslashq−m
to split the integrand into two pieces that cancel upon shifting the integration variable
p→p+q. Recall from exercise (II.7.2) that we were concerned that in some cases the
shift may not be allowed, but it is allowed if the integral is sufficiently convergent, as isindeed the case now that we have regularized. In any event, the proof is in the eating ofthe pudding, and we will see by explicit calculation that /Pi1
μν(q) indeed has the form in (3).
Having learned various computational tricks in the previous chapter you are now ready
to tackle the calculation. I will help by walking you through it. In order not to clutter upthe page I will suppress the regulator terms in (7) in the intermediate steps and restorethem toward the end. After a few steps you should obtain
i/Pi1μν(q)=−/integraldisplayd4p
(2π)4Nμν
D
where Nμν=tr[γν(/negationslashp+ /negationslashq+m)γμ(/negationslashp+m)] and
1
D=/integraldisplay1
0dα1
D
withD=[l2+α(1−α)q2−m2+iε]2, where l=p+αq. Eliminating pin favor of land
beating on Nμνyou will find that Nμνis effectively equal to
−4/parenleftbigg1
2gμνl2+α(1−α)(2qμqν−gμνq2)−m2gμν/parenrightbigg
Integrate over lusing (D.12) and (D.13) and, writing the contribution from the regulators
explicitly, obtain
/Pi1μν(q)=−1
4π2/integraldisplay1
0dα/bracketleftBigg
Fμν(m)−/summationdisplay
acaFμν(ma)/bracketrightBigg
(10)
where
Fμν(m)
=1
2gμν/braceleftbigg
/Lambda12−2[m2−α(1−α)q2] log/Lambda12
m2−α(1−α)q2+m2−α(1−α)q2/bracerightbigg
−[α(1−α)(2qμqν−gμνq2)−m2gμν]/bracketleftbigg
log/Lambda12
m2−α(1−α)q2−1/bracketrightbigg
(11)
Remember, you are doing the calculation; I am just pointing the way. In appendix D, /Lambda1was
introduced to give meaning to various divergent integrals. Since our integral is convergent,
204 | III. Renormalization and Gauge Invariance
we should not need /Lambda1, and indeed, it is gratifying to see that in (10) /Lambda1drops out thanks
to the conditions (8) and (9). Some other terms drop out as well, and we end up with
/Pi1μν(q)=−1
2π2(qμqν−gμνq2)/integraldisplay1
0dα α( 1−α)
{log[m2−α(1−α)q2]−/summationdisplay
acalog[m2
a−α(1−α)q2]} (12)
Lo and behold! The vacuum polarization tensor indeed has the form /Pi1μν(q)=(qμqν−
gμνq2)/Pi1(q2). Our regularization scheme does respect gauge invariance.
Forq2/lessmuchm2
a(the kinematic regime we are interested in had better be much lower than
our threshold of ignorance) we simply define log M2≡/summationtext
acalogm2
ain (12) and obtain
/Pi1(q2)=1
2π2/integraldisplay1
0dα α( 1−α)logM2
m2−α(1−α)q2(13)
Note that our heuristic argument is indeed correct. In the end, effectively we need only
one regulator, but in the intermediate steps we needed two. Actually, this bickering overthe number of regulators is beside the point.
In chapter III.1 I mentioned dimensional regularization as an alternative to Pauli-Villars
regularization. Historically, dimensional regularization was invented to preserve gaugeinvariance in nonabelian gauge theories (which I will discuss in a later chapter). It isinstructive to calculate /Pi1using dimensional regularization (exercise III.7.1).
Electric charge
Physically, we end up with a result for /Pi1(q2)containing a parameter M2expressing our
threshold of ignorance. We conclude that
e2
R=e2 1
1+(e2/12π2)log(M2/m2)/similarequale2/parenleftbigg
1−e2
12π2logM2
m2/parenrightbigg
(14)
Quantum fluctuations effectively diminish the charge. I will explain the physical origin of
this effect in a later chapter on renormalization group flow.
You might argue that physically charge is measured by how strongly one electron scatters
off another electron. T o order e4, in addition to the diagrams in chapter II.7, we also have,
among others, the diagrams shown in figure III.7.4a,b,c. We have computed 4a, but whatabout 4b and 4c? In many texts, it is shown that contributions of III.7.4b and III.7.4c tocharge renormalization cancel. The advantage of using the Lagrangian in (1) is that thisfact becomes self-evident: Charge is a measure of how the photon propagates.
T o belabor a more or less self-evident point let us imagine doing physical or renormalized
perturbation theory as explained in chapter III.3. The Lagrangian is written in terms ofphysical or renormalized fields (and as before we drop the subscript Pon the fields)
L=¯ψ(iγμ(∂μ−iAμ)−mP)ψ−1
4e2
PFμνFμν
+A¯ψiγμ(∂μ−iAμ)ψ+B¯ψψ−CFμνFμν(15)
III.7. Polarizing the Vacuum | 205
+
(a) (b) (c)
Figure III.7.4
where the coefficients of the counterterms A,B, and Care determined iteratively. The
point is that gauge invariance guarantees that ¯ψiγμ∂μψand¯ψγμAμψalways occur in
the combination ¯ψiγμ(∂μ−iAμ)ψ: The strength of the coupling of Aμto¯ψγμψcannot
change. What can change is the ease with which the photon propagates through spacetime.
This statement has profound physical implications. Experimentally, it is known to a
high degree of accuracy that the charges of the electron and the proton are opposite andexactly equal. If the charges were not exactly equal, there would be a residual electrostaticforce between macroscopic objects. Suppose we discovered a principle that tells us thatthe bare charges of the electron and the proton are exactly equal (indeed, as we will see,in grand unification theories, this fact follows from group theory). How do we knowthat quantum fluctuations would not make the charges slightly unequal? After all, theproton participates in the strong interaction and the electron does not and thus manymore diagrams would contribute to the long range electromagnetic scattering betweentwo protons. The discussion here makes clear that this equality will be maintained for theobvious reason that charge renormalization has to do with the photon. In the end, it is alldue to gauge invariance.
Modifying the Coulomb potential
We have focused on charge renormalization, which is determined completely by /Pi1(0), but
in (13) we obtained the complete function /Pi1(q2), which tells us how the qdependence of
the photon propagator is modified. According to the discussion in chapter I.5, the Coulombpotential is just the Fourier transform of the photon propagator (see also exercise III.5.3).Thus, the Coulomb interaction is modified from the venerable 1 /rlaw at a distance scale
of the order of (2m)
−1, namely the inverse of the characteristic value of qin/Pi1(q2). This
modification was experimentally verified as part of the Lamb shift in atomic spectroscopy,another great triumph of quantum electrodynamics.
206 | III. Renormalization and Gauge Invariance
Exercises
III.7.1 Calculate /Pi1μν(q) using dimensional regularization. The procedure is to start with (6), evaluate the trace
inNμν, shift the integration momentum from ptol, and so forth, proceeding exactly as in the text,
until you have to integrate over the loop momentum l. At that point you “pretend” that you are living
ind-dimensional spacetime, so that the term like lμlνinNμν, for example, is to be effectively replaced
by(1/d)g μνl2. The integration is to be performed using (III.1.15) and various generalizations thereof.
Show that the form (3) automatically emerges when you continue to d=4.
III.7.2 Study the modified Coulomb’s law as determined by the Fourier integral/integraltext
d3q{1//vectorq2[1+e2/Pi1(/vectorq2)]}ei/vectorq/vectorx.
III.8 Becoming Imaginary and Conserving Probability
When Feynman amplitudes go imaginary
Let us admire the polarized vacuum, viz (III.7.13):
/Pi1(q2)=1
2π2/integraldisplay1
0dαα( 1−α)log/Lambda12
m2−α(1−α)q2−iε(1)
Dear reader, you have come a long way in quantum field theory, to be able to calculate such
an amazing effect. Quantum fluctuations alter the way a photon propagates!
For a spacelike photon, with q2negative, /Pi1is real and positive for momentum small
compared to the threshold of our ignorance /Lambda1. For a timelike photon, we see that if q2>0
is large enough, the argument of the logarithm may go negative, and thus /Pi1becomes
complex. As you know, the logarithmic function log zcould be defined in the complex
zplane with a cut that can be taken conventionally to go along the negative real axis,
so that for wreal and positive, log (−w±iε)=log(w)±iπ[since in polar coordinates
log(ρeiθ)=log(ρ)+iθ].
We now invite ourselves to define a function in the complex plane: /Pi1(z)≡
1/(2π2)/integraltext1
0dαα( 1−α)log/Lambda12/(m2−α(1−α)z) . The integrand has a cut on the positive
realzaxis extending from z=m2/(α(1−α)) to infinity (fig. III.8.1). Since the maximum
value of α(1−α)in the integration range is1
4,/Pi1(z) is an analytic function in the complex
zplane with a cut along the real axis starting at zc=4m2. The integral over αsmears all
those cuts of the integrand into one single cut.
For timelike photons with large enough q2, a mathematician might be paralyzed won-
dering which side of the cut to go to, but we as physicists know, as per the iεprescription
from chapter I.3, that we should approach the cut from above, namely that we should take/Pi1(q
2+iε) (withε, as always, a positive infinitesimal) as the physical value. Ultimately,
causality tells us which side of the cut we should be on.
That the imaginary part of /Pi1starts at/radicalbig
q2>2mprovides a strong hint of the physics
behind amplitudes going complex. We began the preceding chapter talking about how
208 | III. Renormalization and Gauge Invariance
z
4m2
Figure III.8.1
a photon merrily propagating along could always metamorphose into a virtual electron-
positron pair that, after a short time dictated by the uncertainly principle, annihilate eachother to become a photon again. For/radicalbig
q2>2mthe pair is no longer condemned to be
virtual and to fluctuate out of existence almost immediately. The pair has enough energyto get real. (If you did the exercises religiously, you would recognize that these points werealready developed in exercises I.7.4 and III.1.2.)
Physically, we could argue more forcefully as follows. Imagine a gauge boson of mass M
coupling to electrons just like the photon. (Indeed, in this book we started out supposingthat the photon has a mass.) The vacuum polarization diagram then provides a one-loopcorrection to the vector boson propagator. For M> 2m, the vector boson becomes unstable
against decay into an electron-positron pair. At the same time, /Pi1acquires an imaginary
part. You might suspect that Im /Pi1might have something to do with the decay rate. We will
verify these suspicions later and show that, hey, your physical intuition is pretty good.
When we ended the preceding chapter talking about the modifications to the Coulomb
potential, we thought of a spacelike virtual photon being exchanged between two charges asin electron-electron or electron-proton scattering (chapter II.6). Use crossing (chapter II.8)to map electron-electron scattering into electron-positron scattering. The vacuum polariza-tion diagram then appears (fig. III.8.2) as a correction to electron-positron scattering. Onefunction /Pi1covers two different physical situations.
Incidentally, the title of this section should, strictly speaking, have the word “complex,”
but it is more dramatic to say “When Feynman amplitudes go imaginary,” if only to echocertain movie titles.
Dispersion relations and high frequency behavior
One of the most remarkable discoveries in elementary
particle physics has been that of the existence of thecomplex plane.
—J. Schwinger
Considering that amplitudes are calculated in quantum field theory as integrals over
products of propagators, it is more or less clear that amplitudes are analytic functions
III.8. Becoming Imaginary | 209
e/H11001e/H11002
e/H11001e/H11002
Figure III.8.2
of the external kinematic variables. Another example is the scattering amplitude Min
chapter III.1: it is manifestly an analytic function of s,t, anduwith various cuts. From the
late 1950s until the early 1960s, considerable effort was devoted to studying analyticity inquantum field theory, resulting in a vast literature.
Here we merely touch upon some elementary aspects. Let us start with an embarrass-
ingly simple baby example: f( z)=/integraltext
1
0dα1/(z−α)=log((z−1)/z). The integrand has a
pole at z=α, which got smeared by the integral over αinto a cut stretching from 0 to 1.
At the level of physicist rigor, we may think of a cut as a lot of poles mashed together anda pole as an infinitesimally short cut.
We will mostly encounter real analytic functions, namely functions satisfying f
∗(z)=
f( z∗)(such as log z). Furthermore, we focus on functions that have cuts along the real axis,
as exemplified by /Pi1(z) . For the class of analytic functions specified here, the discontinuity
of the function across the cut is given by disc f( x)≡f( x+iε)−f( x−iε)=f( x+iε)−
f( x+iε)∗=2iImf( x+iε). Define, for σreal,ρ(σ)=Imf( σ+iε). Using Cauchy’s
theorem with a contour Cthat goes around the cut as indicated in figure III.8.3, we could
write
f( z)=/contintegraldisplay
Cdz/prime
2πif( z/prime)
z/prime−z. (2)
Assuming that f( z) vanishes faster than 1 /zasz→∞ , we can drop the contribution from
infinity and write
f( z)=1
π/integraldisplay
dσρ(σ)
σ−z(3)
where the integral ranges over the cut. Note that we can check this equation using the
identity (I.2.14):
Imf( x+iε)=1
π/integraldisplay
dσρ(σ) Im1
σ−x−iε=1
π/integraldisplay
dσρ(σ)πδ(σ −x)
210 | III. Renormalization and Gauge Invariance
z
C
Figure III.8.3
This relation tells us that knowing the imaginary part of falong the cut allows us to
construct fin its entirety, including a fortiori its real part on and away from the cut.
Relations of this type, known collectively as dispersion relations, go back at least to thework of Kramers and Kr ¨onig on optics and are enormously useful in many areas of physics.
We will use it in, for example, chapter VII.4.
We implicitly assumed that the integral over σconverges. If not, we can always (formally)
subtract f(0)=
1
π/integraltext
dσρ(σ)/σ fromf( z) as given above and write
f( z)=f(0)+z
π/integraldisplay
dσρ(σ)
σ(σ−z)(4)
The integral over σnow enjoys an additional factor of 1 /σand hence is more convergent.
In this case, to reconstruct f( z) , we need, in addition to knowledge of the imaginary part
offon the cut, an unknown constant f(0). Evidently, we could repeat this process until
we obtain a convergent integral.
A bell rings, and you, the astute reader, see the connection with the renormalization
procedure of introducing counterterms. In the dispersion weltanschauung, divergentFeynman integrals correspond to integrals over σthat do not converge. Once again,
divergent integrals do not bend real physicists out of shape: we simply admit to ignoranceof the high σregime.
During the height of the dispersion program, it was jokingly said that particle theorists
either group or disperse, depending on whether you like group theory or complex analysisbetter.
III.8. Becoming Imaginary | 211
Imaginary part of Feynman integrals
Going back to the calculation of vacuum polarization in the preceding chapter, we see
that the numerator Nμν, which comes from the spin of the photon and of the electron,
is irrelevant in determining the analytic structure of the Feynman diagram. It is thedenominator Dthat counts. Thus, to get at a conceptual understanding of analyticity
in quantum field theory, we could dispense with spins and study the analog of vacuumpolarization in the scalar field theory with the interaction term L=g(η
†ξ†ϕ+h.c.),
introduced in the appendix to chapter II.6. The ϕpropagator is corrected by the analog
of the diagrams in figure III.7.1 to [compare with (4)]
iDP(q)=i
q2−M2+i/epsilon1+i
q2−M2+i/epsilon1i/Pi1(q2)i
q2−M2+i/epsilon1+...
=i
q2−M2+/Pi1(q2)+i/epsilon1(5)
T o order g2we have
i/Pi1(q2)=i4g2/integraldisplayd4k
(2π)41
k2−μ2+iε1
(q−k)2−m2+iε(6)
As in the preceding chapter we need to regulate the integral, but we will leave that implicit.
Having practiced with the spinful calculation of the preceding chapter, you can now
whiz through this spinless calculation and obtain
/Pi1(z)=g2
16π2/integraldisplay1
0dα log/Lambda12
αm2+(1−α)μ2−α(1−α)z(7)
with/Lambda1some cutoff. (Please do whiz and not imagine that you could whiz.) We use the
same Greek letter /Pi1and allow the two particles in the loop to have different masses, in
contrast to the situation in quantum electrodynamics.
As before, for zreal and negative, the argument of the log is real and positive, and /Pi1is
real. By the same token, for zreal and positive enough, the argument of the log becomes
negative for some value of α, and/Pi1(z) goes complex. Indeed,
Im/Pi1(σ+iε)=−g2
16π2/integraldisplay1
0dα(−π)θ[ α(1−α)σ−αm2−(1−α)μ2]
=g2
16π/integraldisplayα+
α−dα
=g2
16πσ/radicalbig
(σ−(m+μ)2)(σ−(m−μ)2) (8)
withα±the two roots of the quadratic equation obtained by setting the argument of the
step function to zero.
212 | III. Renormalization and Gauge Invariance
Decay and distintegration
At this point you might already be flipping back to the expression given in chapter II.6 for
the decay rate of a particle. Earlier we entertained the suspicion that the imaginary part of/Pi1(z) corresponds to decay. T o confirm our suspicion, let us first go back to elementary
quantum mechanics. The higher energy levels in a hydrogen atom, say, are, strictlyspeaking, not eigenstates of the Hamiltonian: an electron in a higher energy level will, in afinite time, emit a photon and jump to a lower energy level. Phenomenologically, however,the level could be assigned a complex energy E−i
1
2/Gamma1. The probability of staying in this
level then goes with time like |ψ(t)|2∝|e−i(E−i1
2/Gamma1)t|2=e−/Gamma1t. (Note that in elementary
quantum mechanics, the Coulomb and radiation components of the electromagnetic fieldare treated separately: the former is included in the Schr ¨odinger equation but not the latter.
One of the aims of quantum field theory is to remedy this artificial split.)
We now go back to (5) and field theory: note that /Pi1(q
2)effectively shifts M2→M2−
/Pi1(q2). Recall from (III.3.3) that we have counter terms available to, well, counter two
cutoff-dependent pieces of /Pi1(q2). But we have nothing to counter the imaginary part of
/Pi1(q2)with, and so it better be cutoff independent. Indeed it is! The cutoff only appears in
the real part in (7).
We conclude that the effect of /Pi1going imaginary is to shift the mass of the ϕmeson
by a cutoff-independent amount from Mto/radicalbig
M2−iIm/Pi1(M2)≈M−iIm/Pi1(M2)/(2M).
Note that to order g2it suffices to evaluate /Pi1at the unshifted mass squared M2, since
the shift in mass is itself of order g2. Thus /Gamma1=Im/Pi1(M2)/M gives the decay rate, as we
suspected. We obtain (g has dimension of mass and so the dimension is correct)
/Gamma1=g2
16πM3/radicalBig
[M2−(m+μ)2][M2−(m−μ)2] (9)
precisely what we had in (II.6.7). You and I could both take a bow for getting all the factors
exactly right!
Note that both the treatment given in elementary quantum mechanics and here are in
the spirit of treating the decay as a small perturbation. As the width becomes large, at somepoint it no longer makes good sense to talk of the field associated with the particle ϕ.
Taking the imaginary part directly
We ought to be able to take the imaginary part of the Feynman integral in (6) directly, rather
than having to first calculate it as an integral over the Feynman parameter α. I will now
show you how to do this using a trick. For clarity and convenience, change notation from(6), label the momentum carried by the two internal lines in figure III.8.4a separately, andrestore the momentum conservation delta function, so that
i/Pi1(q) =(ig)2i2/integraldisplayd4kη
(2π)4d4kξ
(2π)4(2π)4δ4(kη+kξ−q)/braceleftBigg
1
k2
η−m2
η+i/epsilon11
k2
ξ−m2
ξ+i/epsilon1/bracerightBigg
(10)
III.8. Becoming Imaginary | 213
qk/H9264k/H9257
/H9272ox
/H9264/H9257
(b) (a)
Figure III.8.4
Write the propagator as 1 /(k2−m2+i/epsilon1)=P(1/(k2−m2))−iπδ(k2−m2)and, noting
an explicit overall factor of i, take the real part of the curly bracket above, thus obtaining
Im/Pi1(q)=−g2/integraldisplay
d/Phi1(PηPξ−/Delta1η/Delta1ξ) (11)
For the sake of compactness, we have introduced the notation
d/Phi1=d4kη
(2π)4d4kξ
(2π)4(2π)4δ4(kη+kξ−q),Pη=P1
k2
η−m2
η,/Delta1η=πδ(k2
η−m2
η)
and so on.
We welcome the product of two delta functions; they are what we want, restricting the
two particles ηandξon shell. But yuck, what do we do with the product of the two principal
values? They don’t correspond to anything too physical that we know of.
T o get rid of the two principal values, we use a trick.1First, we regress and recall that
we started out with Feynman diagrams as spacetime diagrams (for example, fig. I.7.6) ofthe process under study. Here (fig. III.8.4b) a ϕexcitation turns into an ηand a ξwith
amplitude igat some spacetime point, which by translation invariance we could take to be
the origin; the ηand the ξexcitations propagate to some point xwith amplitude iD
η(x)
andiDξ(x), respectively, and then recombine into ϕwith amplitude ig(note: not −ig ).
Fourier transforming this product of spacetime amplitudes gives
i/Pi1(q) =(ig)2i2/integraldisplay
dxe−iqxDη(x)D ξ(x) (12)
T o see that this is indeed the same as (10), all you have to do is to plug in the expression
(I.3.22) for Dη(x) andDξ(x).
Incidentally, while many “professors of Feynman diagrams” think almost exclusively
in momentum space, Feynman titled his 1949 paper “Space-Time Approach to QuantumElectrodynamics,” and on occasions it is useful to think of the spacetime roots of a givenFeynman diagram. Now is one of those occasions.
1C. Itzkyson and J.-B. Zuber, Quantum Field Theory, p. 367.
214 | III. Renormalization and Gauge Invariance
Next, go back to exercise I.3.3 and recall that the advanced propagator Dadv(x) and
retarded propagator Dret(x)vanish for x0<0 andx0>0, respectively, and thus the product
Dadv(k)D ret(k) manifestly vanishes for all x. Also recall that the advanced and retarded
propagators Dadv(k)andDret(k)differ from the Feynman propagator D(k) by simply, but
crucially, having their poles in different half-planes in the complex k0plane. Thus
0=−ig2/integraldisplay
dxe−iqxDη, adv(x)D ξ, ret(x)
=/integraldisplayd4kη
(2π)4d4kξ
(2π)4(2π)4δ4(kη+kξ−q)1
k2
η−m2
η−iση/epsilon11
k2
ξ−m2
ξ+iσξ/epsilon1(13)
where we used the shorthand ση=sgn(k0
η)andσξ=sgn(k0
ξ). (The sign function is defined
by sgn (x)=± 1 according to whether x> 0o r<0.) T aking the imaginary part of 0, we
obtain [compare (11)]
0=−g2/integraldisplay
d/Phi1(PηPξ+σησξ/Delta1η/Delta1ξ) (14)
Subtracting (14) from (11) to get rid of the rather unpleasant term PηPξ, we find finally
Im/Pi1(q)=+g2/integraldisplay
d/Phi1( 1+σησξ)(/Delta1η/Delta1ξ)
=g2π2/integraldisplayd4kη
(2π)4d4kξ
(2π)4(2π)4δ4(kη+kξ−q)θ(k0
η)δ(k2
η−m2
η)θ(k0
ξ)δ(k2
ξ−m2
ξ)(1+σησξ)
(15)
Thus/Pi1(q) develops an imaginary part only when the three delta functions can be satisfied
simultaneously.
T o see what these three conditions imply, we can, since /Pi1(q) is a function of q2,g o
to a frame in which q=(Q,/vector0)withQ>0 with no loss of generality. Since k0
η+k0
ξ=
Q>0 and since (1+σησξ)vanishes unless k0
ηandk0
ξhave the same sign, k0
ηandk0
ξ
must be both positive if Im /Pi1(q) is to be nonzero, but that is already mandated by the
two step functions. Furthermore, we need to solve the conservation of energy condition
Q=/radicalBig
/vectork2+m2
η+/radicalBig
/vectork2+m2
ξfor some 3-vector /vectork. This is possible only if Q>m η+mξ,i n
which case, using the identity (I.8.14)
θ(k0)δ(k2−μ2)=θ(k0)δ(k0−εk)
2εk, (16)
we obtain
Im/Pi1(q)=1
2g2/integraldisplayd3kη
(2π)32ωηd3kξ
(2π)32ωξ(2π)4δ4(kη+kξ−q) (17)
We see that (and as we will see more generally) Im /Pi1(q) works out to be a finite integral
over delta functions. Indeed, no counter term is needed.
Some readers might feel that this trick of invoking the advanced and retarded propa-
gators is perhaps a bit “too tricky.” For them, I will show a more brute force method inappendix 1.
III.8. Becoming Imaginary | 215
Unitarity and the Cutkosky cutting rule
The simple example we just went through in detail illustrates what is known as the
Cutkosky cutting rule, which states that to calculate the imaginary part of a Feynmanamplitude we first “cut” through a diagram (as indicated by the dotted line in figure II.8.4a).For each internal line cut, replace the propagator 1 /(k
2−m2+i/epsilon1)byδ(k2−m2); in other
words, put the virtual excitation propagating through the cut onto the mass shell. Thus,in our example, we could jump from (10) to (15) directly. This validates our intuition thatFeynman amplitudes go imaginary when virtual particles can “get real.”
For a precise statement of the cutting rule, see below. (The Cutkosky cut is not be
confused with the Cauchy cut in the complex plane, of course.)
The Cutkosky cutting rule in fact follows in all generality from unitarity. A basic postulate
of quantum mechanics is that the time evolution operator e
−iHTis unitary and hence
preserves probability. Recall from chapter I.8 that it is convenient to split off from theS matrix S
fi=/angbracketleftf|e−iHT|i/angbracketrightthe piece corresponding to “nothing is happening”: S=I+
iT. Unitarity S†S=Ithen implies 2 Im T=i(T†−T)=T†T. Sandwiching this between
initial and final states and inserting a complete set of intermediate states (1 =/summationtext
n|n/angbracketright/angbracketleftn| )
we have
2I mTfi=/summationdisplay
nT†
fnTni (18)
which some readers might recognize as a generalization of the optical theorem from
elementary quantum mechanics.
It is convenient to introduce F=−iM. (We are merely taking out an explicit factor of
iinM: in our simple example, Mcorresponds to i/Pi1,Fto/Pi1.) Then the relation (I.8.16)
between TandMbecomes Tfi=(2π)4δ(4)(Pfi)(/Pi1fi1/ρ)F(f←i), where for the sake
of compactness we have introduced some obvious notations [thus (/Pi1fi1
ρ)denotes the
product of the normalization factors 1 /ρ(see chapter I.8), one for each of the particle in
the state iand in the state f, andPfithe sum of the momenta in fminus the sum of the
momenta in i.]
With this notation, the left-hand side of the generalized optical theorem becomes
2ImTfi=2(2π)4δ(4)(Pfi)(/Pi1fi1/ρ)Im F(f←i)and the right-hand side
/summationdisplay
nT†
fnTni=/summationdisplay
n(2π)4δ(4)(Pfn)(2π)4δ(4)(Pni)/parenleftbigg
/Pi1fn1
ρ/parenrightbigg/parenleftbigg
/Pi1ni1
ρ/parenrightbigg
(F(n←f) )∗F(n←i)
The product of two delta functions δ(4)(Pfn)δ(4)(Pni)=δ(4)(Pfi)δ(4)(Pni), and thus we
could cancel off δ(4)(Pfi). Also (/Pi1fn1/ρ)(/Pi1 ni1/ρ)/(/Pi1 fi1/ρ)=(/Pi1n1/ρ2), and we happily
recover the more familiar factor ρ2[namely (2π)32ωfor bosons]. Thus finally, the gener-
alized optical theorem tells that
2ImF(f←i)=/summationdisplay
n(2π)4δ(4)(Pni)/parenleftbigg
/Pi1n1
ρ2/parenrightbigg
(F(n←f) )∗F(n←i), (19)
216 | III. Renormalization and Gauge Invariance
namely, that the imaginary part of the Feynman amplitude F(f←i)is given by a sum
of(F(n←f) )∗F(n←i)over intermediate states |n/angbracketright. The particles in the intermediate
state are of course physical and on shell. We are to sum over all possible states |n/angbracketrightallowed
by quantum numbers and by the kinematics.
According to Cutkosky, given a Feynman diagram, to obtain its imaginary part, we
simply cut through it in the several different ways allowed, corresponding to the differentpossible intermediate state |n/angbracketright. Particles in the state |n/angbracketrightare manifestly real, not virtual.
Note that unitarity and hence the optical theorem are nonlinear in the transition am-
plitude. This has proved to be enormously useful in actual computation. Suppose we areperturbing in some coupling gand we know F(n←i)andF(n←f)to order g
N. The op-
tical theorem gives us Im F(f←i)to order gN+1, and we could then construct F(f←i)
to order gN+1using a dispersion relation.
The application of the Cutkosky rule to the vacuum polarization function discussed in
this chapter is particularly simple: there is only one possible way of cutting the Feynmandiagram for /Pi1. Here the initial and final states |i/angbracketrightand|f/angbracketrightboth consist of a single ϕmeson,
while the intermediate state |n> consists of an ηand a ξmeson.
Referring to the appendix to chapter II.6, we recall that/summationtext
ncorresponds to/integraltextd3kη
(2π)3d3kξ
(2π)3.
Thus the optical theorem as stated in (19) says that
Im/Pi1(q)=1
2g2/integraldisplayd3kη
(2π)32ωηd3kξ
(2π)32ωξ(2π)4δ4(kη+kξ−q) (20)
precisely what we obtained in (17). You and I could take another bow, since we even get
the factor of 2 correctly (as we must!).
We should also say, in concluding this chapter, “Vive Cauchy!”
Appendix 1: Taking the imaginary part by brute force
For those readers who like brute force, we will extract the imaginary part of
/Pi1=−ig2/integraldisplayd4k
(2π)41
k2−μ2+iε1
(q−k)2−m2+iε(21)
by a more staightforward method, as promised in the text. Since /Pi1depends only on q2, we have the luxury
of setting q=(M,/vector0). We already know that in the complex M2plane, /Pi1has a cut on the axis starting at
M2=(m+μ)2. Let us verify this by brute force.
We could restrict ourselves to M> 0. Factorizing, we find that the denominator of the integrand is a product of
four factors, k0−(εk−iε),k0+(εk−iε),k0−(M+Ek−iε), andk0−(M−Ek+iε), and thus the integrand
has four poles in the complex k0plane. (Evidently, εk=/radicalBig
/vectork2+μ2,Ek=/radicalbig
/vectork2+m2, and if you are perplexed over
the difference between εkandε, then you are hopelessly confused.) We now integrate over k0, choosing to close
the contour in the lower half plane and going around picking up poles. Picking up the pole at εk−iε, we obtain
/Pi11=−g2/integraltext
(d3k/(2π)3)(1/(2εk(εk−M−Ek)(εk−M+Ek))). Picking up the pole at M+Ek−iε, we obtain
/Pi12=−g2/integraltext
(d3k/(2π)3)(1/((M+Ek−εk)(M+Ek+εk)(2Ek))). We now regard /Pi1=/Pi11+/Pi12as a function of
M:
/Pi1=−g2/integraldisplayd3k
(2π)31
(M+Ek−εk)/bracketleftbigg1
2εk(M−εk−Ek+iε)+1
2Ek(M+Ek+εk−iε)/bracketrightbigg
(22)
III.8. Becoming Imaginary | 217
In spite of appearances, there is no pole at M≈εk−Ek. (Since this pole would lead to a cut at μ−m, there better
not be!) For M> 0 we only care about the pole at M≈εk+Ek=/radicalBig
/vectork2+μ2+/radicalbig
/vectork2+m2. When we integrate over
/vectork, this pole gets smeared into a cut starting at m+μ. So far so good.
T o calculate the discontinuity across the cut, we use the identity (16) once again Restoring the iε’s and throwing
away the term we don’t care about, we have effectively
/Pi1=−g2/integraldisplayd3k
(2π)31
2εk(εk−M−Ek+iε)(ε k−M+Ek−iε)
=− 2πg2/integraldisplayd4k
(2π)4θ(k0)δ(k2−μ2)1
(M−εk+Ek−iε)(M −(εk+Ek)+iε)
The discontinuity of /Pi1across the cut just specified is determined by applying (I.2.13) to the factor
1/(M−(εk+Ek)+iε), giving Im /Pi1=2π2g2/integraltext
(d4k/(2π)4)θ(k0)δ(k2−μ2)δ(M−(εk+Ek))/2Ek. Use the
identity (16) again in the form
θ(q0−k0)θ((q−k)2−m2)=θ(q0−k0)θ((q0−k0)−Ek)
2Ek(23)
and we obtain
Im/Pi1=2π2g2/integraldisplayd4k
(2π)4θ(k0)δ(k2−μ2)θ(q0−k0)δ((q−k)2−m2) (24)
Remarkably, as Cutkosky taught us, to obtain the imaginary part we simply replace the propagators in (21) by
delta functions.
Appendix 2: A dispersion representation for the two-point amplitude
I would like to give you a bit more flavor of the dispersion program once active and now being revived (seepart N). Consider the two-point amplitude iD(x)≡/angbracketleft0|T(O(x)O(0))|0/angbracketright, with O(x)some operator in the canonical
formalism. For example, for O(x) equal to the field ϕ(x) ,D(x) would be the propagator. In chapter I.8, we were
able to evaluate D(x) for a free field theory, because then we could solve the field equation of motion and expand
ϕ(x) in terms of creation and annihilation operators. But what can we do in a fully interacting field theory? There
is no hope of solving the operator field equation of motion.
The goal of the dispersion program of the 1950s and 1960s is to say as much as possible about D(x) based on
general considerations such as analyticity.
OK, so first write iD(x)=θ(x
0)/angbracketleft0|eiPxO(0)e−iPxO(0)|0/angbracketright+θ( −x0)/angbracketleft0|O(0)eiPxO(0)e−iPx|0/angbracketright, where we
used spacetime translation O(x)=eiP.xO(0)e−iP.x. By the way, if you are not totally sure of this relation, dif-
ferentiate it to obtain ∂μO(x)=i[Pμ,O(x)] which you should recognize as the relativistic version of the usual
Heisenberg equations (I.8.2, 3). Now insert 1 =/summationtext
n|n/angbracketright/angbracketleftn| , with |n/angbracketrighta complete set of intermediate states, to obtain
/angbracketleft0|eiPxO(0)e−iPxO(0)|0/angbracketright=/angbracketleft 0|O(0)e−iPxO(0)|0/angbracketright=/summationtext
n/angbracketleft0|O(0)|n/angbracketright/angbracketleftn| e−iPxO(0)|0/angbracketright=/summationtext
ne−iPnx|O0n|2, where
we used Pμ|0/angbracketright=0 and Pμ|n/angbracketright=Pμ
n|n/angbracketright and defined O0n≡/angbracketleft0|O(0)|n/angbracketright. Next, use the integral representations
for the step function θ(t)=−i/integraltext
(dω/2 π)eiωt/(ω−iε)andθ(−t)=i/integraltext
(dω/2 π)eiωt/(ω+iε). Again, if you are
not sure of this, simply differentiated
dtθ(t)=−id
dt/integraltext
(dω/2 π)eiωt/(ω−iε)=/integraltext
(dω/2 π)eiωt, which you recog-
nize from (I.2.12) as indeed the integral representation of the delta function δ(t)=d
dtθ(t) . In other words, the
representation used here is the integral of the representation in (I.2.12).
Putting it all together, we obtain
iD(q)=/integraldisplay
d4xeiq.xiD(x)=−i(2π)3/summationdisplay
n|O0n|2/braceleftBigg
δ(3)(/vectorq−/vectorPn)
P0
n−q0−iε+δ(3)(/vectorq+/vectorPn)
P0
n+q0−iε/bracerightBigg
. (25)
The integral over d3xproduced the 3-dimensional delta function, while the integral over dx0=dtpicked up the
denominator in the integral representation for the step function.
Now take the imaginary part using Im1 /(P0
n−q0−iε)=πδ(q0−P0
n). We thus obtain
Im(i/integraldisplay
d4xeiqx/angbracketleft0|T(O(x)O(0))|0/angbracketright)=π(2π)3/summationdisplay
n|O0n|2(δ(4)(q−Pn)+δ(4)(q+Pn)) (26)
218 | III. Renormalization and Gauge Invariance
with the more pleasing 4-dimensional delta function. Note that for q0>0 the term involving δ(4)(q+Pn)drops
out, since the energies of physical states must be positive.
What have we accomplished? Even though we are totally incapable of calculating D(q), we have managed to
represent its imaginary part in terms of physical quantities that are measurable in principle, namely the absolutesquare |O
0n|2of the matrix element of O(0)between the vacuum state and the state |n/angbracketright. For example, if O(x) is
the meson field ϕ(x) in aϕ4theory, the state |n/angbracketrightwould consist of the single-meson state, the three-meson state,
and so on. The general hope during the dispersion era was that by keeping a few states we could obtain a decentapproximation to D(q). Note that the result does not depend on perturbing in some coupling constant.
The contribution of the single meson state |/vectork/angbracketrighthas a particularly simple form, as you might expect.
With our normalization of single-particle states (as in chapter I.8), Lorentz invariance implies /angbracketleft/vectork|O(0)|0/angbracketright=
Z1
2//radicalbig
(2π)32ωk, with ωk=/radicalbig
/vectork2+m2andZ1
2an unknown constant, measuring the “strength” with which Ois
capable of producing the single meson from the vacuum. Putting this into (25) and recognizing that the sum
over single-meson states is now given by/integraltext
d3k|/vectork/angbracketright/angbracketleft/vectork|[with the normalization /angbracketleft/vectork/prime|/vectork/angbracketright=δ3(/vectork/prime−/vectork)], we find that
the single-meson contribution to iD(q) is given by
−i( 2π)3/integraldisplay
d3kZ
(2π)32ωk/braceleftBigg
δ(3)(/vectorq−/vectork)
ωk−q0−iε+(q→−q)/bracerightBigg
=−iZ
2ωq/braceleftBigg
1
ωq−q0−iε+(q→−q)/bracerightBigg
=iZ
q2−m2+iε(27)
This is a very satisfying result: even though we cannot calculate iD(q), we know that it has a pole at a position
determined by the meson mass with a residue that depends on how Ois normalized.
As a check, we can also easily calculate the contribution of the single-meson state to −ImD(q). Plugging
into (26), we find, for q0>0,πZ/integraltext
(d3k/2ωk)δ4(q−k)=(πZ/2 ωq)δ(q0−ωq)=πZδ(q2−m2), where we used
(I.8.14, 16) in the last step.
Given our experience with the vacuum polarization function, we would expect D(q) (which by Lorentz
invariance is a function of q2) to have a cut starting at q2=(3m)2. T o verify this, simply look at (26) and choose /vectorq=
0. The contribution of the three-meson state occurs at/radicalbig
q2=q0=P0
“3”=/radicalBig
/vectork2
1+m2+/radicalBig
/vectork2
2+m2+/radicalBig
/vectork2
3+m2≥
3m. The sum over states is now a triple integration over /vectork1,/vectork2, and/vectork3, subject to the constraint /vectork1+/vectork2+/vectork3=0.
Knowing the imaginary part of D(q) we can now write a dispersion relation of the kind in (3).
Finally, if you stare at (26) long enough (see exercise III.8.3) you will discover the relation
Im(i/integraldisplay
d4xeiqx/angbracketleft0|T(O(x)O(0))|0/angbracketright)=1
2/integraldisplay
d4xeiqx/angbracketleft0|[O(x),O(0)]|0/angbracketright (28)
The discussion here is relevant to the discussion of field redefinition in chapter I.8. Suppose our friend uses
η=Z1
2ϕ+αϕ3instead of ϕ; then the present discussion shows that his propagator/integraltext
d4xeiqx/angbracketleft0|T(η(x)η( 0))|0/angbracketright
still has a pole at q2−m2. The important point is that physics fixes the pole to be at the same location.
Here we have taken Oto be a Lorentz scalar. In applications (see chapter VII.3) the role of Ois often played
by the electromagnetic current Jμ(x) (treated as an operator). The same discussion holds except that we have
to keep track of some Lorentz indices. Indeed, we recognize that the vacuum polarization function /Pi1μνthen
corresponds to the function Din this discussion.
Exercises
III.8.1 Evaluate the imaginary part of the vacuum polarization function, and by explicit calculation verify that it
is related to the decay rate of a vector particle into an electron and a positron.
III.8.2 Suppose we add a term gϕ3to our scalar ϕ4theory. Show that to order g4there is a “box diagram”
contributing to meson scattering p1+p2→p3+p4with the amplitude
I=g4/integraldisplayd4k
(2π)41
(k2−m2−iε)((k +p2)2−m2−iε)((k −p1)2−m2−iε)((k +p2−p3)2−m2−iε)
III.8. Becoming Imaginary | 219
Calculate the integral explicitly as a function of s=(p1+p2)2andt=(p3−p2)2. Study the analyticity
property of Ias a function of sfor fixed t. Evaluate the discontinuity of Iacross the cut and verify
Cutkosky’s cutting rule. Check that the optical theorem works. What about the analyticity property of I
as a function of tfor fixed s? And as a function of u=(p3−p1)2?
III.8.3 Prove (28). [Hint: Do unto/integraltext
d4xeiqx/angbracketleft0|[O(x),O(0)]|0/angbracketrightwhat we did to/integraltext
d4xeiqx/angbracketleft0||T(O(x)O(0))|0/angbracketright,
namely, insert 1 =/summationtext
n|n/angbracketright/angbracketleftn| (with |n/angbracketrighta complete set of states) between O(x) andO(0)in the commu-
tator. Now we don’t have to bother with representing the step function.]
This page intentionally left blank
Part IV Symmetry and Symmetry Breaking
This page intentionally left blank
IV.1 Symmetry Breaking
A symmetric world would be dull
While we would like to believe that the fundamental laws of Nature are symmetric,
a completely symmetric world would be rather dull, and as a matter of fact, the realworld is not perfectly symmetric. More precisely, we want the Lagrangian, but not theworld described by the Lagrangian, to be symmetric. Indeed, a central theme of modernphysics is the study of how symmetries of the Lagrangian can be broken. We will see insubsequent chapters that our present understanding of the fundamental laws is built uponan understanding of symmetry breaking.
Consider the Lagrangian studied in chapter I.10:
L=1
2/bracketleftBig
(∂/vectorϕ)2−μ2/vectorϕ2/bracketrightBig
−λ
4(/vectorϕ2)2(1)
where /vectorϕ=(ϕ1,ϕ2,... ,ϕN). This Lagrangian exhibits an O(N) symmetry under which /vectorϕ
transforms as an N-component vector.
We can easily add terms that do not respect the symmetry. For instance, add terms
such as ϕ2
1,ϕ4
1andϕ2
1/vectorϕ2and break the O(N) symmetry down to O(N−1), under which
ϕ2,..., ϕNrotate as an (N −1)-component vector. This way of breaking the symmetry,
“by hand” as it were, is known as explicit breaking.
We can break the symmetry in stages. Obviously, if we want to, we can break it down to
O(N−M) by hand, for any M<N .
Note that in this example, with the terms we added, the reflection symmetry ϕa→−ϕa
(anya)still holds. It is easy enough to break this symmetry as well, by adding a term such
asϕ3
a, for example.
Breaking the symmetry by hand is not very interesting. Indeed, we might as well start
with a nonsymmetric Lagrangian in the first place.
224 | IV . Symmetry and Symmetry Breaking
V(q)
q
Figure IV .1.1
Spontaneous symmetry breaking
A more subtle and interesting way is to let the system “break the symmetry itself,” a
phenomenon known as spontaneous symmetry breaking. I will explain by way of anexample. Let us flip the sign of the /vectorϕ
2term in (1) and write
L=1
2/bracketleftBig
(∂/vectorϕ)2+μ2/vectorϕ2/bracketrightBig
−λ
4(/vectorϕ2)2(2)
Naively, we would conclude that for small λthe field ϕcreates a particle of mass/radicalbig
−μ2=
iμ. Something is obviously wrong.
The essential physics is analogous to what would happen if we give the spring constant
in an anharmonic oscillator the wrong sign and write L=1
2(˙q2+kq2)−(λ/4)q4. We all
know what to do in classical mechanics. The potential energy V( q)=−1
2kq2+(λ/4)q4
[known as the double-well potential (figure IV .1.1)] has two minima at q=±v, where
v≡(k/λ)1
2. At low energies, we choose either one of the two minima and study small
oscillations around that minimum. Committing to one or the other of the two minimabreaks the reflection symmetry q→−qof the system.
In quantum mechanics, however, the particle can tunnel between the two minima, the
tunneling barrier being V(0)−V(±v) . The probability of being in one or the other of
the two minima must be equal, thus respecting the reflection symmetry q→−qof the
Hamiltonian. In particular, the ground state wave function ψ(q)=ψ(−q) is even.
Let us try to extend the same reasoning to quantum field theory. For a generic scalar
field Lagrangian L=
1
2(∂0ϕ)2−1
2(∂iϕ)2−V( ϕ) we again have to find the minimum of the
potential energy/integraltext
dDx[1
2(∂iϕ)2+V( ϕ) ], where Dis the dimension of space. Clearly, any
spatial variation in ϕonly increases the energy, and so we set ϕ(x) to equal a spacetime
independent quantity ϕand look for the minimum of V( ϕ) . In particular, for the example
in (2), we have
V( ϕ)=−1
2μ2/vectorϕ2+λ
4(/vectorϕ2)2(3)
As we will see, the N=1 case is dramatically different from the N≥2 cases.
IV .1. Symmetry Breaking | 225
Difference between quantum mechanics and quantum field theory
Study the N=1 case first. The potential V( ϕ) looks exactly the same as the potential in
figure IV .1.1 with the horizontal axis relabeled as ϕ. There are two minima at ϕ=±v=
±(μ2/λ)1
2.
But some thought reveals a crucial difference between quantum field theory and quan-
tum mechanics. The tunneling barrier is now [ V(0)−V(±v) ]/integraltext
dDx(where Ddenotes
the dimension of space) and hence infinite (or more precisely, extensive with the volume ofthe system)! T unneling is shut down, and the ground state wave function is concentratedaround either +v or−v. We have to commit to one or the other of the two possibilities
for the ground state and build perturbation theory around it. It does not matter whichone we choose: The physics is equivalent. But by making a choice, we break the reflectionsymmetry ϕ→−ϕof the Lagrangian.
The reflection symmetry is broken spontaneously! We did not put symmetry breaking
terms into the Lagrangian by hand but yet the reflection symmetry is broken.
1
Let’s choose the ground state at +v and write ϕ=v+ϕ/prime. Expanding in ϕ/primewe find after
a bit of arithmetic that
L=μ4
4λ+1
2(∂ϕ/prime)2−μ2ϕ/prime2−O(ϕ/prime3) (4)
The physical particle created by the shifted field ϕ/primehas mass√
2μ. The physical mass
squared has to come out positive since, after all, it is just −V/prime/prime(ϕ)|ϕ=v, as you can see after
a moment’s thought.
Similarly, you would recognize that the first term in (4) is just −V( ϕ) |ϕ=v. If we are only
interested in the scattering of the mesons associated with ϕ/primethis term does not enter at all.
Indeed, we are always free to add an arbitrary constant to Lto begin with. We had quite
arbitrarily set V( ϕ=0)equal to 0. The same situation appears in quantum mechanics: In
the discussion of the harmonic oscillator the zero point energy1
2/planckover2piωis not observable; only
transitions between energy levels are physical. We will return to this point in chapter VIII.2.
Yet another way of looking at (2) is that quantum field theory amounts to doing the
Euclidean functional integral
Z=/integraldisplay
Dϕe−/integraltext
ddx{1
2[(∂ϕ)2−μ2ϕ2]+λ
4(ϕ2)2}
and perturbation theory just corresponds to studying the small oscillations around a
minimum of the Euclidean action. Normally, with μ2positive, we expand around the
minimum ϕ=0. With μ2negative, ϕ=0 is a local maximum and not a minimum.
In quantum field theory what is called the ground state is also known as the vacuum,
since it is literally the state in which the field is “at rest,” with no particles present. Here we
1An insignificant technical aside for the nitpickers: Strictly speaking, in field theory the ground state wave
function should be called a wave functional, since /Psi1[ϕ(/vectorx)] is a functional of the function ϕ(/vectorx).
226 | IV . Symmetry and Symmetry Breaking
have two physically equivalent vacua from which we are to choose one. The value assumed
byϕin the ground state, either vor−vin our example, is known as the vacuum expectation
value of ϕ. The field ϕis said to have acquired a vacuum expectation value.
Continuous symmetry
Let us now turn to (2) with N≥2. The potential (3) is shown in figure IV .1.2 for N=2. The
shape of the potential has been variously compared to the bottom of a punted wine bottleor a Mexican hat. The potential is minimized at /vectorϕ
2=μ2/λ. Something interesting is going
on: We have an infinite number of vacua characterized by the direction of /vectorϕin that vacuum.
Because of the O(2)symmetry of the Lagrangian they are all physically equivalent. The
result had better not depend on our choice. So let us choose /vectorϕto point in the 1 direction,
that is, ϕ1=v≡+/radicalbig
μ2/λandϕ2=0.
Now consider fluctuations around this field configuration, in other words, write ϕ1=
v+ϕ/prime
1andϕ2=ϕ/prime
2, plug into (2) for N=2, and expand Lout. I invite you to do the
arithmetic. You should find (after dropping the primes on the fields; why clutter thenotation, right?)
L=μ4
4λ+1
2/bracketleftBig
(∂ϕ 1)2+(∂ϕ2)2/bracketrightBig
−μ2ϕ2
1+O(ϕ3) (5)
The constant term is exactly as in (4), and just like the field ϕ/primein (4), the field ϕ1has mass√
2μ. But now note the remarkable feature of (5): the absence of a ϕ2
2term. The field ϕ2is
massless!
Emergence of massless boson
Thatϕ2comes out massless is not an accident. I will now explain that the masslessness is
a general and exact phenomenon.
Referring back to figure IV .1.2 we can easily understand the particle spectrum. Excitation
in the ϕ1field corresponds to fluctuation in the radial direction, “climbing the wall” so to
speak, while excitation in the ϕ2field corresponds to fluctuation in the angular direction,
“rolling along the gutter” so to speak. It costs no energy for a marble to roll along theminima of the potential energy, going from one minimum to another. Another way ofsaying this is to picture a long wavelength excitation of the form ϕ
2=asin(ωt−/vectork/vectorx)with
asmall. In a region of length scale small compared to |/vectork|−1, the field ϕ2is essentially
constant and thus the field /vectorϕis just rotated slightly away from the 1 direction, which by
theO(2)symmetry is equivalent to the vacuum. It is only when we look at regions of
length scale large compared to |/vectork|−1that we realize that the excitation costs energy. Thus,
as|/vectork|→0, we expect the energy of the excitation to vanish.
We now understand the crucial difference between the N=1 and the N=2 cases: In
the former we have a reflection symmetry, which is discrete, while in the latter we have anO(2)symmetry, which is continuous.
IV .1. Symmetry Breaking | 227
Figure IV .1.2
We have worked out the N=2 case in detail. You should now be able to generalize our
discussion to arbitrary N≥2 (see exercise IV .1.1).
Meanwhile, it is worth looking at N=2 from another point of view. Many field theories
can be written in more than one form and it is important to know them under differentguises. Construct the complex field ϕ=(1/√
2)(ϕ1+iϕ2); we have ϕ†ϕ=1
2(ϕ2
1+ϕ2
2)and
so can write (2) as
L=∂ϕ†∂ϕ+μ2ϕ†ϕ−λ(ϕ†ϕ)2(6)
which is manifestly invariant under the U(1)transformation ϕ→eiαϕ(recall chapter I.10).
You may recognize that this amounts to saying that the groups O(2)andU(1)are locally
isomorphic. Just as we can write a vector in Cartesian or polar coordinates we are freeto parametrize the field by ϕ(x)=ρ(x)e
iθ(x)(as in chapter III.5) so that ∂μϕ=(∂μρ+
iρ∂μθ)eiθ. We obtain L=ρ2(∂θ)2+(∂ρ)2+μ2ρ2−λρ4. Spontaneous symmetry breaking
means setting ρ=v+χwithv=+/radicalbig
μ2/2λ, whereupon
L=v2(∂θ)2+⎡
⎣(∂χ)2−2μ2χ2−4/radicalBigg
μ2λ
2χ3−λχ4⎤
⎦+⎛
⎝/radicalBigg
2μ2
λχ+χ2⎞
⎠(∂θ)2(7)
We recognize the phase θ(x) as the massless field. We have arranged the terms in the
Lagrangian in three groups: the kinetic energy of the massless field θ, the kinetic and
potential energy of the massive field χ, and the interaction between θandχ. (The additive
constant in (5) has been dropped to minimize clutter.)
228 | IV . Symmetry and Symmetry Breaking
Goldstone’s theorem
We will now prove Goldstone’s theorem, which states that whenever a continuous sym-
metry is spontaneously broken, massless fields, known as Nambu2-Goldstone bosons,
emerge.
Recall that associated with every continuous symmetry is a conserved charge Q. That Q
generates a symmetry is stated as
[H,Q]=0 (8)
Let the vacuum (or ground state in quantum mechanics) be denoted by |0/angbracketright. By adding
an appropriate constant to the Hamiltonian H→H+cwe can always write H|0/angbracketright=0.
Normally, the vacuum is invariant under the symmetry transformation, eiθQ|0/angbracketright=| 0/angbracketright,o r
in other words Q|0/angbracketright=0.
But suppose the symmetry is spontaneously broken, so that the vacuum is not invariant
under the symmetry transformation; in other words, Q|0/angbracketright /negationslash= 0. Consider the state Q|0/angbracketright.
What is its energy? Well,
HQ|0/angbracketright=[H,Q]|0/angbracketright=0 (9)
[The first equality follows from H|0/angbracketright=0 and the second from (8).] Thus, we have found
another state Q|0/angbracketrightwith the same energy as |0/angbracketright.
Note that the proof makes no reference to either relativity or fields. You can also see that
it merely formalizes the picture of the marble rolling along the gutter.
In quantum field theory, we have local currents, and so
Q=/integraldisplay
dDxJ0(/vectorx,t)
where Ddenotes the dimension of space and conservation of Qsays that the integral can
be evaluated at any time. Consider the state
|s/angbracketright=/integraldisplay
dDxe−i/vectork/vectorxJ0(/vectorx,t)|0/angbracketright
which has3spatial momentum /vectork.A s/vectorkgoes to zero it goes over to Q|0/angbracketright, which as we learned
in (9) has zero energy. Thus, as the momentum of the state |s/angbracketrightgoes to zero, its energy goes
to zero. In a relativistic theory, this means precisely that |s/angbracketrightdescribes a massless particle.
2Y . Nambu, quite deservedly, received the 2008 physics Nobel Prize for his profound contribution to our
understanding of spontaneous symmetry breaking.
3Acting on it with Pi(exercise I.11.3) and using Pi|0/angbracketright=0, we have
Pi|s/angbracketright=/integraldisplay
dDxe−i/vectork/vectorx[Pi,J0(/vectorx,t)]|0/angbracketright=−i/integraldisplay
dDe−i/vectork./vectorx∂iJ0(/vectorx,t)|0/angbracketright=ki|s/angbracketright
upon integrating by parts.
IV .1. Symmetry Breaking | 229
The proof makes clear that the theorem practically exudes generality: It applies to any
spontaneously broken continuous symmetry.
Counting Nambu-Goldstone bosons
From our proof, we see that the number of Nambu-Goldstone bosons is clearly equal tothe number of conserved charges that do not leave the vacuum invariant, that is, do notannihilate |0/angbracketright. For each such charge Q
α, we can construct a zero-energy state Qα|0/angbracketright.
In our example, we have only one current Jμ=i(ϕ1∂μϕ2−ϕ2∂μϕ1)and hence one
Nambu-Goldstone boson. In general, if the Lagrangian is left invariant by a symmetrygroup Gwithn(G) generators, but the vacuum is left invariant by only a subgroup HofG
withn(H) generators, then there are n(G)− n(H) Nambu-Goldstone bosons. If you want
to show off your mastery of mathematical jargon you can say that the Nambu-Goldstonebosons live in the coset space G/H .
Ferromagnet and spin wave
The generality of the proof suggests that the usefulness of Goldstone’s theorem is not
restricted to particle physics. In fact, it originated in condensed matter physics, the classicexample there being the ferromagnet. The Hamiltonian, being composed of just theinteraction of nonrelativistic electrons with the ions in the solid, is of course invariantunder the rotation group SO( 3), but the magnetization /vectorMpicks out a direction, and
the ferromagnet is left invariant only under the subgroup SO( 2)consisting of rotations
about the axis defined by /vectorM. The Nambu-Goldstone theorem is easy to visualize physically.
Consider a “spin wave” in which the local magnetization /vectorM(/vectorx)varies slowly from point
to point. A physicist living in a region small compared to the wavelength does not evenrealize that he or she is no longer in the “vacuum.” Thus, the frequency of the wave mustgo to zero as the wavelength goes to infinity. This is of course exactly the same heuristicargument given earlier. Note that quantum mechanics is needed only to translate the wavevector /vectorkinto momentum and the frequency ωinto energy. I will come back to magnets
and spin wave in chapters V .3 and VI.5.
Quantum fluctuations and the dimension of spacetime
Our discussion of spontaneous symmetry breaking is essentially classical. What happenswhen quantum fluctuations are included? I will address this question in detail in chap-ter IV .3, but for now let us go back to (5). In the ground state, ϕ
1=vandϕ2=0. Recall
that in the mattress model of a scalar field theory the mass term comes from the springsholding the mattress to its equilibrium position. The term −μ
2ϕ/prime2
1(note the prime) in (5)
230 | IV . Symmetry and Symmetry Breaking
tells us that it costs action for ϕ/prime
1to wander away from its ground state value ϕ/prime
1=0. But
now we are worried: ϕ2is massless. Can it wander away from its ground state value? T o
answer this question let us calculate the mean square fluctuation
/angbracketleft(ϕ2(0))2/angbracketright=1
Z/integraldisplay
DϕeiS(ϕ)[ϕ2(0)]2
=lim
x→01
Z/integraldisplay
DϕeiS(ϕ)ϕ2(x)ϕ 2(0)
=lim
x→0/integraldisplayddk
(2π)deikx
k2(10)
(We recognized the functional integral that defines the propagator; recall chapter I.7.)
The upper limit of the integral in (10) is cut off at some /Lambda1(which would correspond to
the inverse of the lattice spacing when applying these ideas to a ferromagnet) and so asexplained in chapter III.1 (and as you will see in chapter VIII.3) we are not particularlyworried about the ultraviolet divergence for large k. But we do have to worry about a possible
infrared divergence for small k.(Note that for a massive field 1 /k
2in (10) would have been
replaced by 1 /(k2+μ2)and there would be no infrared divergence.)
We see that there is no infrared divergence for d> 2. Our picture of spontaneously
breaking a continuous symmetry is valid in our (3+1)-dimensional world.
However, for d≤2 the mean square fluctuation of ϕ2comes out infinite, so our naive
picture is totally off. We have arrived at the Coleman-Mermin-Wagner theorem (provedindependently by a particle theorist and two condensed matter theorists), which statesthat spontaneous breaking of a continuous symmetry is impossible for d=2. Note that
while our discussion is given for O(2)symmetry the conclusion applies to any continuous
symmetry since the argument depends only on the presence of Nambu-Goldstone fields.
In our examples, symmetry is spontaneously broken by a scalar field ϕ, but nothing says
that the field ϕmust be elementary. In many condensed matter systems, superconductors,
for example, symmetries are spontaneously broken, but we know that the system consistsof electrons and atomic nuclei. The field ϕis generated dynamically, for example as a bound
state of two electrons in superconductors. More on this in chapter V .4. The spontaneousbreaking of a symmetry by a dynamically generated field is sometimes referred to asdynamical symmetry breaking.
4
Exercises
IV .1.1 Show explicitly that there are N−1 Nambu-Goldstone bosons in the G=O(N) example (2).
IV .1.2 Construct the analog of (2) with Ncomplex scalar fields and invariant under SU(N) . Count the number
of Nambu-Goldstone bosons when one of the scalar fields acquires a vacuum expectation value.
4This chapter is dedicated to the memory of the late Jorge Swieca.
IV.2 The Pion as a Nambu-Goldstone Boson
Crisis for field theory
After the spectacular triumphs of quantum field theory in the electromagnetic interaction,
physicists in the 1950s and 1960s were naturally eager to apply it to the strong and weakinteractions. As we have already seen, field theory when applied to the weak interactionappeared not to be renormalizable. As for the strong interaction, field theory appeared to-tally untenable for other reasons. For one thing, as the number of experimentally observedhadrons (namely strongly interacting particles) proliferated, it became clear that were weto associate a field with each hadron the resulting field theory would be quite a mess, withnumerous arbitrary coupling constants. But even if we were to restrict ourselves to nucle-ons and pions, the known coupling constant of the interaction between pions and nucleonsis a large number. (Hence the term strong interaction in the first place!) The perturbativeapproach that worked so spectacularly well in quantum electrodynamics was doomed tofailure.
Many eminent physicists at the time advocated abandoning quantum field theory al-
together, and at certain graduate schools, quantum field theory was even dropped fromthe curriculum. It was not until the early 1970s that quantum field theory made a tri-umphant comeback. A field theory for the strong interaction was formulated, not in termsof hadrons, but in terms of quarks and gluons. I will get to that in chapter VII.3.
Pion weak decay
T o understand the crisis facing field theory, let us go back in time and imagine what a fieldtheorist might be trying to do in the late 1950s. Since this is not a book on particle physics,I will merely sketch the relevant facts. You are urged to consult one of the texts on the
subject.
1By that time, many semileptonic decays such as n→p+e−+ν,π−→e−+ν,
1See, e.g., E. Commins and P. H. Bucksbaum, Weak Interactions of Leptons and Quarks.
232 | IV . Symmetry and Symmetry Breaking
andπ−→π0+e−+νhad been measured. Neutron βdecay n→p+e−+νwas of
course the process for which Fermi invented his theory, which by that time had assumedthe form L=G[
eγμ(1−γ5)ν][pγμ(1−γ5)n], where nis a neutron field annihilating a
neutron, pa proton field annihilating a proton, νa neutrino field annihilating a neutrino
(or creating an antineutrino as in βdecay), and ean electron field annihilating an electron.
It became clear that to write down a field for each hadron and a Lagrangian for each
decay process, as theorists were in fact doing for a while, was a losing battle. Instead, weshould write
L=G[eγμ(1−γ5)ν](Jμ−J5μ) (1)
withJμandJ5μtwo currents transforming as a Lorentz vector, and axial vector respectively.
We think of JμandJ5μas quantum operators in a canonical formulation of field theory.
Our task would then be to calculate the matrix elements between hadron states, /angbracketleftp|(Jμ−
J5μ)|n/angbracketright,/angbracketleft0|(Jμ−J5μ)|π−/angbracketright,/angbracketleftπ0|(Jμ−J5μ)|π−/angbracketright, and so on, corresponding to the three
decay processes I listed above. (I should make clear that although I am talking about weakdecays, the calculation of these matrix elements is a problem in the strong interaction. Inother words, in understanding these decays, we have to treat the strong interaction to allorders in the strong coupling, but it suffices to treat the weak interaction to lowest order inthe weak coupling G.) Actually, there is a precedent for the attitude we are adopting here.
T o account for nuclear βdecay (Z,A)→(Z+1,A)+e
−+ν, Fermi certainly did not
write a separate Lagrangian for each nucleus. Rather, it was the task of the nuclear theoristto calculate the matrix element /angbracketleftZ+1,A|[
pγμ(1−γ5)n]|Z,A/angbracketright. Similarly, it is the task of
the strong interaction theorist to calculate matrix elements such as /angbracketleftp|(Jμ−J5μ)|n/angbracketright.
For the story I am telling, let me focus on trying to calculate the matrix element of
the axial vector current Jμ
5between a neutron and a proton. Here we make a trivial
change in notation: We no longer indicate that we have a neutron in the initial state and aproton in the final state, but instead we specify the momentum pof the neutron and the
momentum p
/primeof the proton. Incidentally, in (1) the fields and the currents are of course
all functions of the spacetime coordinates x. Thus, we want to calculate /angbracketleftp/prime|Jμ
5(x)|p/angbracketright,
but by translation invariance this is equal to /angbracketleftp/prime|Jμ
5(0)|p/angbracketrighte−i(p/prime−p).x. Henceforth, we
simply calculate /angbracketleftp/prime|Jμ
5(0)|p/angbracketrightand suppress the 0. Note that spin labels have already been
suppressed.
Lorentz invariance and parity can take us some distance: They imply that2
/angbracketleftp/prime|Jμ
5|p/angbracketright=¯u(p/prime)[γμγ5F(q2)+qμγ5G(q2)]u(p) (2)
withq≡p/prime−p[compare with (III.6.7)]. But Lorentz invariance and parity can only take
us so far: We know nothing about the “form factors” F(q2)andG(q2).
2Another possible term of the form (p/prime+p)μγ5can be shown to vanish by charge conjugation and isospin
symmetries.
IV .2. Pion as Nambu-Goldstone Boson | 233
(a) (b)
Figure IV .2.1
Similarly, for the matrix element /angbracketleft0|Jμ
5|π−/angbracketrightLorentz invariance tells us that
/angbracketleft0|Jμ
5|k/angbracketright=fkμ(3)
I have again labeled the initial state by the momentum kof the pion. The right-hand side
of (3) has to be a vector but since kis the only vector available it has to be proportional
tok. Just like F(q2)andG(q2), the constant fis a strong interaction quantity that we
don’t know how to calculate. On the other hand, F(q2),G(q2), and fcan all be measured
experimentally. For instance, the rate for the decay π−→e−+νclearly depends on f2.
Too many diagrams
Let us look over the shoulder of a field theorist trying to calculate /angbracketleftp/prime|Jμ
5|p/angbracketrightand/angbracketleft0|Jμ
5|k/angbracketright
in (2)in the late 1950s. He would draw Feynman diagrams such as the ones in figures IV .2.1
and IV .2.2 and soon realize that it would be hopeless. Because of the strong coupling, hewould have to calculate an infinite number of diagrams, even if the strong interaction weredescribed by a field theory, a notion already rejected by many luminaries of the time.
Figure IV .2.2
234 | IV . Symmetry and Symmetry Breaking
In telling the story of the breakthrough I am not going to follow the absolutely fascinating
history of the subject, full of total confusion and blind alleys. Instead, with the benefit ofhindsight, I am going to tell the story using what I regard as the best pedagogical approach.
The pion is very light
The breakthrough originated in the observation that the mass of the π−at 139 Mev was
considerably less than the mass of the proton at 938 Mev. For a long time this was simplytaken as a fact not in any particular need of an explanation. But eventually some theoristswondered why one hadron should be so much lighter than another.
Finally, some theorists took the bold step of imagining an “ideal world” in which the
π
−is massless. The idea was that this ideal world would be a good approximation of our
world, to an accuracy of about 15% (∼139/938).
Do you remember one circumstance in which a massless spinless particle would emerge
naturally? Yes, spontaneous symmetry breaking! In one of the blinding insights that havecharacterized the history of particle physics, some theorists proposed that the πmesons
are the Nambu-Goldstone bosons of some spontaneous broken symmetry.
Indeed, let’s multiply (3) by k
μ:
kμ/angbracketleft0|Jμ
5|k/angbracketright=fk2=fm2
π(4)
which is equal to zero in the ideal world. Recall from our earlier discussion on translation
invariance that
/angbracketleft0|Jμ
5(x)|k/angbracketright=/angbracketleft 0|Jμ
5(0)|k/angbracketrighte−ik.x
and hence
/angbracketleft0|∂μJμ
5(x)|k/angbracketright=−ikμ/angbracketleft0|Jμ
5(0)|k/angbracketrighte−ik.x
Thus, if the axial current is conserved, ∂μJμ
5(x)=0, in the ideal world, kμ/angbracketleft0|Jμ
5|k/angbracketright=0
and (4) would indeed imply m2
π=0.
The ideal world we are discussing enjoys a symmetry known as the chiral symmetry
of the strong interaction. The symmetry is spontaneously broken in the ground state weinhabit, with the πmeson as the Nambu-Goldstone boson. The Noether current associated
with this symmetry is the conserved J
μ
5.
In fact, you should recognize that the manipulation here is closely related to the proof
of the Nambu-Goldstone theorem given in chapter IV .1.
Goldberger-Treiman relation
Now comes the punchline. Multiply (2 )by(p/prime−p)μ. By the same translation invariance
argument we just used,
(p/prime−p)μ/angbracketleftp/prime|Jμ
5(0)|p/angbracketright=i/angbracketleftp/prime|∂μJμ
5(x)|p/angbracketrightei(p/prime−p).x
IV .2. Pion as Nambu-Goldstone Boson | 235
and hence vanishes if ∂μJμ
5=0. On the other hand, multiplying the right-hand side of
(2)by(p/prime−p)μwe obtain ¯u(p/prime)[(/negationslashp/prime− /negationslashp)γ5F(q2)+q2γ5G(q2)]u(p) . Using the Dirac
equation (do it!) we conclude that
0=2mNF(q2)+q2G(q2) (5)
withmNthe nucleon mass.
The form factors F(q2)andG(q2)are each determined by an infinite number of
Feynman diagrams we have no hope of calculating, but yet we have managed to relatethem! This represents a common strategy in many areas of physics: When faced withvarious quantities we don’t know how to calculate, we can nevertheless try to relate them.
We can go farther by letting q→0 in (5). Referring to (2) we see that F(0)is measured
experimentally in n→p+e
−+ν(the momentum transfer is negligible on the scale of
the strong interaction). But oops, we seem to have a problem: We predict the nucleon massm
N=0!
In fact, we are saved by examining figure IV .2.1b: There are an infinite number of
diagrams exhibiting a pole due to none other than the massless πmeson, which you can
see gives
fqμ1
q2gπNN¯u(p/prime)γ5u(p) (6)
When the πpropagator joins onto the nucleon line, an infinite number of diagrams
summed together gives the experimentally measured pion-nucleon coupling constantg
πNN . Thus, referring to (2), we see that for q∼0 the form factor G(q2)∼f(1/q2)gπNN .
Plugging into (5), we obtain the celebrated Goldberger-T reiman relation
2mNF(0)+fgπNN=0 (7)
relating four experimentally measured quantities. As might be expected, it holds with about
a 15% error, consistent with our not living in a world with an exactly massless πmeson.
Toward a theory of the strong interaction
The art of relating infinite sets of Feynman diagrams without calculating them, and it is an
art form involving a great deal of cleverness, was developed into a subject called dispersionrelations and S-matrix theory, which we mentioned briefly in chapter III.8. Our present
understanding of the strong interaction was built on this foundation. You could see fromthis example that an important component of dispersion relations was the study of theanalyticity properties of Feynman diagrams as described in chapter III.8. The essence ofthe Goldberger-T reiman argument is separating the infinite number of diagrams into thosewith a pole in the complex q
2-plane and those without a pole (but with a cut.)
The discovery that the strong interaction contains a spontaneously broken symmetry
provided a crucial clue to the underlying theory of the strong interaction and ultimatelyled to the concepts of quarks and gluons.
236 | IV . Symmetry and Symmetry Breaking
A note for the historian of science: Whether theoretical physicists regard a quantity as
small or large depends (obviously) on the cultural and mental framework they grew upin. T reiman once told me that the notion of setting 138 Mev to zero, when the energyreleased per nucleon in nuclear fission is of order 10 Mev, struck the generation that grewup with the atomic bomb (as T reiman did—he was with the armed forces in the Pacific) assurely the height of absurdity. Now of course a new generation of young string theoristsis perfectly comfortable in regarding anything less than the Planck energy 10
19Gev as
essentially zero.
IV.3 Effective Potential
Quantum fluctuations and symmetry breaking
The important phenomenon of spontaneous symmetry breaking was based on minimizing
the classical potential energy V( ϕ) of a quantum field theory. It is natural to wonder how
quantum fluctuations would change this picture.
T o motivate the discussion, consider once again (III.3.3)
L=1
2(∂ϕ)2−1
2μ2ϕ2−1
4!λϕ4+A(∂ϕ)2+Bϕ2+Cϕ4(1)
(Speaking of quantum fluctuations, we have to include counterterms as indicated.) What
have you learned about this theory? For μ2>0, the action is extremized at ϕ=0, and
quantizing the small fluctuations around ϕ=0 we obtain scalar particles that scatter off
each other. For μ2<0, the action is extremized at some ϕmin, and the discrete symmetry
ϕ→−ϕis spontaneously broken, as you learned in chapter IV .1. What happens when
μ=0? T o break or not to break, that is the question.
A quick guess is that quantum fluctuations would break the symmetry. The μ=0 theory
is posed on the edge of symmetry breaking, and quantum fluctuations ought to push itover the brink. Think of a classical pencil perfectly balanced on its tip. Then “switch on”quantum mechanics.
Wisdom of the son-in-law
Let us follow Schwinger and Jona-Lasinio and develop the formalism that enables us toanswer this question. Consider a scalar field theory defined by
Z=eiW(J)=/integraldisplay
Dϕei[S(ϕ)+Jϕ ](2)
[with the convenient shorthand Jϕ=/integraltext
d4xJ(x)ϕ(x)]. If we can do the functional integral,
we obtain the generating functional W(J) . As explained in chapter I.7, by differentiating
238 | IV . Symmetry and Symmetry Breaking
Wwith respect to the source J(x) repeatedly, we can obtain any Green’s function and
hence any scattering amplitude we want. In particular,
ϕc(x)≡δW
δJ(x)=1
Z/integraldisplay
Dϕei[S(ϕ)+Jϕ ]ϕ(x) (3)
The subscript cis used traditionally to remind us (see appendix 2 in chapter I.8) that in a
canonical formalism ϕc(x) is the expectation value /angbracketleft0|ˆϕ|0/angbracketrightof the quantum operator ˆϕ.I t
is certainly not to be confused with the integration dummy variable ϕin (3). The relation
(3) determines ϕc(x) as a functional of J.
Given a functional WofJwe can perform a Legendre transform to obtain a functional
/Gamma1ofϕc. Legendre transform is just the fancy term for the simple relation
/Gamma1(ϕc)=W(J) −/integraldisplay
d4xJ(x)ϕ c(x) (4)
The relation is simple, but be careful about what it says: It defines a functional of ϕc(x)
through the implicit dependence of Jonϕc. On the right-hand side of (4) Jis to be
eliminated in favor of ϕcby solving (3). We expand the functional /Gamma1(ϕc)in the form
/Gamma1(ϕc)=/integraldisplay
d4x[−V eff(ϕc)+Z(ϕc)(∂ϕc)2+...] (5)
where (...)indicates terms with higher and higher powers of ∂. We will soon see the
wisdom of the notation Veff(ϕc).
The point of the Legendre transform is that the functional derivative of /Gamma1is nice and
simple:
δ/Gamma1(ϕc)
δϕc(y)=/integraldisplay
d4xδJ(x)
δϕc(y)δW(J)
δJ(x)−/integraldisplay
d4xδJ(x)
δϕc(y)ϕc(x)−J(y)
=−J(y) (6)
a relation we can think of as the “dual” of δW(J)/δJ(x) =ϕc(x).
If you vaguely feel that you have seen this sort of manipulation before in your physics
eduction, you are quite right! It was in a course on thermodynamics, where you learnedabout the Legendre transform relating the free energy to the energy: F=E−TS with
Fa function of the temperature TandEa function of the entropy S. Thus Jand
ϕare “conjugate” pairs just like TandS(or even more clearly magnetic field Hand
magnetization M). Convince yourself that this is far more than a mere coincidence.
ForJandϕ
cindependent of xwe see from (5) that the condition (6) reduces to
V/prime
eff(ϕc)=J (7)
This relation makes clear what the effective potential Veff(ϕc)is good for. Let’s ask what
happens when there is no external source J. The answer is immediate: (7) tells us that
V/prime
eff(ϕc)=0, (8)
IV .3. Effective Potential | 239
In other words, the vacuum expectation value of ˆϕin the absence of an external source is
determined by minimizing Veff(ϕc).
First order in quantum fluctuations
All of these formal manipulations are not worth much if we cannot evaluate W(J) .I n
fact, in most cases we can only evaluate eiW(J)=/integraltext
Dϕei[S(ϕ)+Jϕ ]in the steepest descent
approximation (see chapter I.2). Let us turn the crank and find the steepest descent “point”ϕ
s(x), namely the solution of (henceforth I will drop the subscript c as there is little risk
of confusion)
δ[S(ϕ)+/integraltext
d4yJ(y)ϕ(y)]
δϕ(x)/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle
ϕs=0 (9)
or more explicitly,
∂2ϕs(x)+V/prime[ϕs(x)]=J(x) (10)
Write the dummy integration variable in (2) as ϕ=ϕs+/tildewideϕand expand to quadratic order
in/tildewideϕto obtain
Z=e(i//planckover2pi)W(J)=/integraldisplay
Dϕe(i//planckover2pi)[S(ϕ)+Jϕ ]
/similarequale(i//planckover2pi)[S(ϕs)+Jϕ s]/integraldisplay
D/tildewideϕe(i//planckover2pi)/integraltext
d4x1
2[(∂/tildewideϕ)2−V/prime/prime(ϕs)/tildewideϕ2]
=e(i//planckover2pi)[S(ϕs)+Jϕ s]−1
2tr log[∂2+V/prime/prime(ϕs)](11)
We have used (II.5.2) to represent the determinant we get upon integrating over /tildewideϕ. Note
that I have put back Planck’s constant /planckover2pi. Here ϕs, as a solution of (10), is to be regarded as
a function of J.
Now that we have determined
W(J) =[S(ϕs)+Jϕs]+i/planckover2pi
2tr log[∂2+V/prime/prime(ϕs)]+O(/planckover2pi2)
it is straightforward to Legendre transform. I will go painfully slowly here:
ϕ=δW
δJ=δ[S(ϕs)+Jϕs]
δϕsδϕs
δJ+ϕs+O(/planckover2pi)=ϕs+O(/planckover2pi)
T o leading order in /planckover2pi,ϕ(namely the object formerly known as ϕc)is equal to ϕs. Thus,
from (4) we obtain
/Gamma1(ϕ)=S(ϕ)+i/planckover2pi
2tr log[∂2+V/prime/prime(ϕ)]+O(/planckover2pi2) (12)
Nice though this formula looks, in practice it is impossible to evaluate the trace for
arbitrary ϕ(x) : We have to find all the eigenvalues of the operator ∂2+V/prime/prime(ϕ), take their
log, and sum. Our task simplifies drastically if we are content with studying /Gamma1(ϕ) for
240 | IV . Symmetry and Symmetry Breaking
ϕindependent of x, in which case V/prime/prime(ϕ) is a constant and the operator ∂2+V/prime/prime(ϕ) is
translation invariant and easily treated in momentum space:
tr log[∂2+V/prime/prime(ϕ)]=/integraldisplay
d4x/angbracketleftx|log[∂2+V/prime/prime(ϕ)]|x/angbracketright
=/integraldisplay
d4x/integraldisplayd4k
(2π)4/angbracketleftx|k/angbracketright/angbracketleftk|log[∂2+V/prime/prime(ϕ)]|k/angbracketright/angbracketleftk|x/angbracketright
=/integraldisplay
d4x/integraldisplayd4k
(2π)4log[−k2+V/prime/prime(ϕ)] (13)
Referring to (5), we obtain
Veff(ϕ)=V( ϕ)−i/planckover2pi
2/integraldisplayd4k
(2π)4log/bracketleftbiggk2−V/prime/prime(ϕ)
k2/bracketrightbigg
+O(/planckover2pi2) (14)
known as the Coleman-Weinberg effective potential. What we computed is the order /planckover2pi
correction to the classical potential V( ϕ) . Note that we have added a ϕindependent constant
to make the argument of the logarithm dimensionless.
We can give a nice physical interpretation of (14). Let the universe be suffused with
the scalar field ϕ(x) taking on the value ϕ, a background field so to speak. For V( ϕ)=
1
2μ2ϕ2+(1/4!)λϕ4, we have V/prime/prime(ϕ)=μ2+1
2λϕ2≡μ(ϕ)2, which, as the notation μ(ϕ)2
suggests, we recognize as the ϕ-dependent effective mass squared of a scalar particle
propagating in the background field ϕ. The mass squared μ2in the Lagrangian is corrected
by a term1
2λϕ2due to the interaction of the particle with the background field ϕ. Now we
see clearly what (14) tells us: The first term V( ϕ) is the classical energy density contained
in the background ϕ, while the second term is the vacuum energy density of a scalar field
with mass squared equal to V/prime/prime(ϕ) [see (II.5.3) and exercise IV .3.4].
Your renormalization theory at work
The integral in (14) is quadratically divergent, or more correctly, quadratically dependent
on the cutoff. But no sweat, we were instructed to introduce three counterterms (of whichonly two are relevant here since ϕis independent of x). Thus, we actually have
Veff(ϕ)=V( ϕ)+/planckover2pi
2/integraldisplayd4kE
(2π)4log/bracketleftBigg
k2
E+V/prime/prime(ϕ)
k2
E/bracketrightBigg
+Bϕ2+Cϕ4+O(/planckover2pi2) (15)
where we have Wick rotated to a Euclidean integral (see appendix D). Using (D.9) and
integrating up to k2
E=/Lambda12, we obtain (suppressing /planckover2pi)
Veff(ϕ)=V( ϕ)+/Lambda12
32π2V/prime/prime(ϕ)−[V/prime/prime(ϕ)]2
64π2loge1
2/Lambda12
V/prime/prime(ϕ)+Bϕ2+Cϕ4(16)
As expected, since the integrand in (15) goes as 1 /k2
Efor large k2
Ethe integral depends
quadratically and logarithmically on the cutoff /Lambda12.
IV .3. Effective Potential | 241
Watch renormalization theory at work! Since Vis a quartic polynomial in ϕ,V/prime/prime(ϕ) is a
quadratic polynomial and [ V/prime/prime(ϕ)]2a quartic polynomial. Thus, we have just enough coun-
terterms Bϕ2+Cϕ4to absorb the cutoff dependence. This is a particularly transparent
example of how the method of adding counterterms works.
T o see how bad things can happen in a nonrenormalizable theory, suppose in contrast
thatVis a polynomial of degree 6 in ϕ. Then we are allowed to have three counterterms
Bϕ2+Cϕ4+Dϕ6, but that is not enough since [ V/prime/prime(ϕ)]2is now a polynomial of degree
8. This means that we should have started out with Va polynomial of degree 8, but then
[V/prime/prime(ϕ)]2would be a polynomial of degree 12. Clearly, the process escalates into an infinite-
degree polynomial. We see the hallmark of a nonrenormalizable theory: its insatiableappetite for counterterms.
Imposing renormalization conditions
Waking up from the nightmare of an infinite number of counterterms chasing us, let usgo back to the sweetly renormalizable ϕ
4theory. In chapter III.3 we fix the counterterms
by imposing conditions on various scattering amplitudes. Here we would have to fix thecoefficients BandCby imposing two conditions on V
eff(ϕ) at appropriate values of ϕ.W e
are working in field space, so to speak, rather than momentum space, but the conceptualframework is the same.
We could proceed with the general quartic polynomial V( ϕ) , but instead let us try to
answer the motivating question of this chapter: What happens when μ=0, that is, when
V( ϕ)=(1/4!)λϕ
4? The arithmetic is also simpler.
Evaluating (16) we get
Veff(ϕ)=(/Lambda12
64π2λ+B)ϕ2+(1
4!λ+λ2
(16π)2logϕ2
/Lambda12+C)ϕ4+O(λ3)
(after absorbing some ϕ-independent constants into C). We see explicitly that the /Lambda1
dependence can be absorbed into BandC.
We started out with a purely quartic V( ϕ) . Quantum fluctuations generate a quadrati-
cally divergent ϕ2term that we can cancel with the Bcounterterm. What does μ=0 mean?
It means that (d2V/ dϕ2)|ϕ=0vanishes. T o say that we have a μ=0 theory means that we
have to maintain a vanishing renormalized mass squared, defined here as the coefficient
ofϕ2. Thus, we impose our first condition
d2Veff
dϕ2/vextendsingle/vextendsingle/vextendsingle/vextendsingle
ϕ=0=0 (17)
This is a somewhat long-winded way of saying that we want B=−(/Lambda12/64π2)λto this
order.
Similarly, we might think that the second condition would be to set (d4Veff/dϕ4)|ϕ=0
equal to some coupling, but differentiating the ϕ4logϕterm in Vefffour times we are
going to get a term like log ϕ, which is not defined at ϕ=0. We are forced to impose our
242 | IV . Symmetry and Symmetry Breaking
condition on d4Veff/dϕ4not at ϕ=0 but at ϕequal to some arbitrarily chosen mass M.
(Recall that ϕhas the dimension of mass.) Thus, the second condition reads
d4Veff
dϕ4/vextendsingle/vextendsingle/vextendsingle/vextendsingle
ϕ=M=λ(M) (18)
where λ(M) is a coupling manifestly dependent on M.
Plugging
Veff(ϕ)=(1
4!λ+λ2
(16π)2logϕ2
/Lambda12+C)ϕ4+O(λ3)
into (18) we see that λ(M) is equal to λplusO(λ2)corrections, among which is a term like
λ2logM. We can get a clean relation by differentiating λ(M) :
Mdλ(M)
dM=3
16π2λ2+O(λ3)
=3
16π2λ(M)2+O[λ(M)3] (19)
where the second equality is correct to the order indicated. This interesting relation tells us
how the coupling λ(M) depends on the mass scale Mat which it is defined. Recall exercise
III.1.3. We will come back to this relation in chapter VI.7 on the renormalization group.
Meanwhile, let us press on. Using (18) to determine Cand plugging it into Veffwe
obtain
Veff(ϕ)=1
4!λ(M)ϕ4+λ(M)2
(16π)2ϕ4/parenleftbigg
logϕ2
M2−25
6/parenrightbigg
+O[λ(M)3] (20)
You are no longer surprised, I suppose, that Cand the cutoff /Lambda1have both disappeared.
That’s a renormalizable theory for you!
The fact that Veffdoes not depend on the arbitrarily chosen M, namely, M(dV eff/dM) =
0, reproduces (19) to the order indicated.
Breaking by quantum fluctuations
Now we can answer the motivating question: T o break or not to break?
Quantum fluctuations generate a correction to the potential of the form +ϕ4logϕ2,
but log ϕ2is whopping big and negative for small ϕ! The O(/planckover2pi)correction overwhelms
the classical O(/planckover2pi0)potential +ϕ4nearϕ=0. Quantum fluctuations break the discrete
symmetry ϕ→−ϕ.
It is easy enough to determine the minima ±ϕmin ofVeff(ϕ) (which you should plot as
a function of ϕto get a feeling for). But closer inspection shows us that we cannot take the
precise value of ϕmin seriously; Veffhas the form λϕ4(1+λlogϕ+...)suggesting that
the expansion parameter is actually λlogϕrather than λ. [T ry to convince yourself that
(...)starts with (λlogϕ)2.] The minima ϕmin ofVeffclearly occurs when the expansion
parameter is of order unity. In an exercise in chapter IV .7 you will see a clever way of gettingaround this problem.
IV .3. Effective Potential | 243
Fermions
In (11) ϕsplays the role of an external field while /tildewideϕcorresponds to a quantum field we
integrate over. The role of /tildewideϕcan also be played by a fermion field ψ. Consider adding
¯ψ(i/negationslash∂−m−fϕ) ψ to the Lagrangian. In the path integral
Z=/integraldisplay
DϕD ¯ψDψei/integraltext
d4x[1
2(∂ϕ)2−V( ϕ) +¯ψ(i/negationslash∂−m−fϕ) ψ ](21)
we can always choose to integrate over ψfirst, obtaining
Z=/integraldisplay
Dϕei/integraltext
d4x[1
2(∂ϕ)2−V( ϕ) ]+tr log (i/negationslash∂−m−fϕ)(22)
Repeating the steps in (13) we find that the fermion field contributes
VF(ϕ)=+i/integraldisplayd4p
(2π)4tr log/negationslashp−m−fϕ
/negationslashp(23)
toVeff(ϕ). (The trace in (23) is taken over the gamma matrices.) Again from chapter II.5,
we see that physically VF(ϕ) represents the vacuum energy of a fermion with the effective
massm(ϕ)≡m+fϕ.
We can massage the trace of the logarithm using tr log M=log det M(II.5.12) and
cyclically permuting factors in a determinant):
tr log(/negationslashp−a)=tr logγ5(/negationslashp−a)γ5=tr log(− /negationslashp−a)
=1
2tr(log(/negationslashp−a)+log(/negationslashp+a))+1
2tr log(−1)
=1
2tr log(−1)(p2−a2). (24)
Hence,
tr log(/negationslashp−a)
/negationslashp=1
2tr logp2−a2
p2=2 logp2−a2
p2(25)
and so
VF(ϕ)=2i/integraldisplayd4p
(2π)4logp2−m(ϕ)2
p2(26)
Contrast the overall sign with the sign in (14): the difference in sign between fermionic
and bosonic loops was explained in chapter II.5.
Thus, in the end the effective potential generated by the quantum fluctuations has a
pleasing interpretation: It is just the energy density due to the fluctuating energy, entirelyanalogous to the zero point energy of the harmonic oscillator, of quantum fields living inthe background ϕ(see exercise IV .3.5).
Exercises
IV .3.1 Consider the effective potential in (0+1)-dimensional spacetime:
Veff(ϕ)=V( ϕ)+/planckover2pi
2/integraldisplaydkE
(2π)logk2
E+V/prime/prime(ϕ)
k2
E+O(/planckover2pi2)
244 | IV . Symmetry and Symmetry Breaking
No counterterm is needed since the integral is perfectly convergent. But (0+1)-dimensional field theory
is just quantum mechanics. Evaluate the integral and show that Veffis in complete accord with your
knowledge of quantum mechanics.
IV .3.2 Study Veffin(1+1)−dimensional spacetime.
IV .3.3 Consider a massless fermion field ψcoupled to a scalar field ϕbyfϕ¯ψψ in(1+1)-dimensional
spacetime. Show that
VF=1
2π(f ϕ)2logϕ2
M2(27)
after a suitable counterterm has been added. This result is important in condensed matter physics, as
we will see in chapter V .5 on the Peierls instability.
IV .3.4 Understand (14) using Feynman diagrams. Show that Veffis generated by an infinite number of di-
agrams. [Hint: Expand the logarithm in (14) as a series in V/prime/prime(ϕ)/k2and try to associate a Feynman
diagram with each term in the series.]
IV .3.5 Consider the electrodynamics of a complex scalar field
L=−1
4FμνFμν+/bracketleftBig
(∂μ+ieAμ)ϕ†/bracketrightBig/bracketleftbig
(∂μ−ieAμ)ϕ/bracketrightbig
+μ2ϕ†ϕ−λ(ϕ†ϕ)2(28)
In a universe suffused with the scalar field ϕ(x) taking on the value ϕindependent of xas in the text,
the Lagrangian will contain a term (e2ϕ†ϕ)AμAμso that the effective mass squared of the photon field
becomes M(ϕ)2≡e2ϕ†ϕ. Show that its contribution to Veff(ϕ) has the form
/integraldisplayd4k
(2π)4logk2−M(ϕ)2
k2(29)
Compare with (14) and (26). [Hint: Use the Landau gauge to simplify the calculation.] If you need help,
I strongly urge you to read S. Coleman and E. Weinberg, Phys. Rev. D7: 1883, 1973, a paragon of clarity
in exposition.
IV.4 Magnetic Monopole
Quantum mechanics and magnetic monopoles
Curiously enough, while electric charges are commonplace nobody has ever seen a mag-
netic charge or monopole. Within classical physics we can perfectly well modify one ofMaxwell’s equations to /vector∇./vectorB=ρ
M, with ρMdenoting the density of magnetic monopoles.
The only price we have to pay is that the magnetic field /vectorBcan no longer be represented
as/vectorB=/vector∇×/vectorAsince otherwise /vector∇./vectorB=/vector∇./vector∇×/vectorA=εijk∂i∂jAk=0 identically. Newton and
Leibniz told us that derivatives commute with each other.
So what, you say. Indeed, who cares that /vectorBcannot be written as /vector∇×/vectorA? The vector po-
tential /vectorAwas introduced into physics only as a mathematical crutch, and indeed that is
still how students are often taught in a course on classical electromagnetism. As the dis-tinguished nineteenth-century physicist Heaviside thundered, “Physics should be purgedof such rubbish as the scalar and vector potentials; only the fields /vectorEand/vectorBare physical.”
With the advent of quantum mechanics, however, Heaviside was proved to be quite
wrong. Recall, for example, the nonrelativistic Schr ¨odinger equation for a charged particle
in an electromagnetic field:
/bracketleftbigg
−1
2m(/vector∇−ie/vectorA)2+eφ/bracketrightbigg
ψ=Eψ (1)
Charged particles couple directly to the vector and scalar potentials /vectorAandφ, which are
thus seen as being more fundamental, in some sense, than the electromagnetic fields /vectorE
and/vectorB, as I alluded to in chapter III.4. Quantum physics demands the vector potential.
Dirac noted brilliantly that these remarks imply an intrinsic conflict between quantum
mechanics and the concept of magnetic monopoles. Upon closer analysis, he found thatquantum mechanics does not actually forbid the existence of magnetic monopoles. Itallows magnetic monopoles, but only those carrying a specific amount of magnetic charge.
246 | IV . Symmetry and Symmetry Breaking
Differential forms
For the following discussion and for the next chapter on Y ang-Mills theory, it is highly
convenient to use the language of differential forms. Fear not, we will need only a fewelementary concepts. Let x
μbeDreal variables (thus, the index μtakes on Dvalues)
andAμ(not necessarily the electromagnetic gauge potential in this purely mathematical
section) be Dfunctions of the x’s. In our applications, xμrepresent coordinates and, as
we will see, differential forms have natural geometric interpretations.
We call the object A≡Aμdxμa 1-form. The differentials dxμare treated following New-
ton and Leibniz. If we change coordinates x→x/prime, then as usual dxμ=(∂xμ/∂x/primeν)dx/primeνso
thatA≡Aμdxμ=Aμ(∂xμ/∂x/primeν)dx/primeν≡A/prime
νdx/primeν. This reproduces the standard transforma-
tion law of vectors under coordinate transformation A/prime
ν=Aμ(∂xμ/∂x/primeν). As an example,
consider A=cosθd ϕ . Regarding θandϕas angular coordinates on a 2-sphere (namely
the surface of a 3-ball), we have Aθ=0 and Aϕ=cosθ. Similarly, we define a p-form as
H=(1/p!)Hμ1μ2...μpdxμ1dxμ2...dxμp. (Repeated indices are summed, as always.) The
“degenerate” example is that of a 0-form, call it /Lambda1, which is just a scalar function of the
coordinates xμ. An example of a 2-form is F=(1/2!)Fμνdxμdxν.
We now face the question of how to think about products of differentials. In an ele-
mentary course on calculus we learned that dx dy represents the area of an infinitesimal
rectangle with length dxand width dy. At that level, we more or less automatically regard
dy dx as the same as dx dy . The order of writing the differentials does not matter. However,
think about making a coordinate transformation so that x=x(x/prime,y/prime)andy=y(x/prime,y/prime)are
now functions of the new coordinates x/primeandy/prime. Now look at
dx dy =/parenleftbigg∂x
∂x/primedx/prime+∂x
∂y/primedy/prime/parenrightbigg/parenleftbigg∂y
∂x/primedx/prime+∂y
∂y/primedy/prime/parenrightbigg
(2)
Note that the coefficient of dx/primedy/primeis(∂x/∂x/prime)(∂y/∂y/prime)and that the coefficient of dy/primedx/prime
is(∂x/∂y/prime)(∂y/∂x/prime). We see that it is much better if we regard the differentials dxμ
as anticommuting objects [what mathematicians would call Grassmann variables (recall
chapter (II.5)] so that dy/primedx/prime=−dx/primedy/primeanddx/primedx/prime=0=dy/primedy/prime. Then (2) simplifies
neatly to
dx dy =/parenleftbigg∂x
∂x/prime∂y
∂y/prime−∂x
∂y/prime∂y
∂x/prime/parenrightbigg
dx/primedy/prime≡J(x ,y;x/prime,y/prime)dx/primedy/prime(3)
We obtain the correct Jacobian J(x ,y;x/prime,y/prime)for transforming the area element dx dy to
the area element dx/primedy/prime.
In many texts, dx dy is written as dx^dy . We will omit the wedge—no reason to clutter
up the page.
This little exercise tells us that we should define dxμdxν=−dxνdxμand regard the
area element dxμdxνas directional. The area elements dxμdxνanddxνdxμhave the same
magnitude but point in opposite directions.
IV .4. Magnetic Monopole | 247
We now define a differential operation dto act on any form. Acting on a p-form H,i t
gives by definition
dH=1
p!∂νHμ1μ2...μpdxνdxμ1dxμ2...dxμp
Thus, d/Lambda1=∂ν/Lambda1dxνand
dA=∂νAμdxνdxμ=1
2(∂νAμ−∂μAν)dxνdxμ
In the last step, we used dxμdxν=−dxνdxμ.
We see that this mathematical formalism is almost tailor made to describe electromag-
netism. If we call A≡Aμdxμthe potential 1-form and think of Aμas the electromag-
netic potential, then F=dA is in fact the field 2-form. If we write Fout in terms of its
components F=(1/2!)Fμνdxμdxν, then Fμνis indeed equal to the electromagnetic field
(∂μAν−∂νAμ).
Note that xμis not a form, and dxμis not dacting on a form.
If you like, you can think of differential forms as “merely” an elegantly compact notation.
The point is to think of physical objects such as AandFas entities, without having to
commit to any particular coordinate system. This is particularly convenient when one hasto deal with objects more complicated than AandF, for example in string theory. By using
differential forms, we avoid drowning in a sea of indices.
An important identity is
dd=0 (4)
which says that acting with don any form twice gives zero. Verify this as an exercise. In
particular dF=ddA=0. If you write this out in components you will recognize it as a
standard identity (the “Bianchi identity”) in electromagnetism.
Closed is not necessarily globally exact
It is convenient here to introduce some jargon. A p-form αis said to be closed if dα=0.
It is said to be exact if there exists a (p−1)-form βsuch that α=dβ.
T alking the talk, we say that (4) tells us that exact forms are closed.Is the converse of (4) true? Kind of. The Poincar ´e lemma states that a closed form is
locally exact. In other words, if dH=0 with Hsome p-form, then locally
H=dK (5)
for some (p−1)-form K. However, it may or may not be the case that H=dK globally,
that is, everywhere. Actually, whether you know it or not, you are already familiar with the
Poincar ´e lemma. For example, surely you learned somewhere that if the curl of a vector
field vanishes, the vector field is locally the gradient of some scalar field.
Forms are ready made to be integrated over. For example, given the 2-form F=
(1/2!)Fμνdxμdxν, we can write/integraltext
MFfor any 2-manifold M. Note that the measure is
248 | IV . Symmetry and Symmetry Breaking
already included and there is no need to specify a coordinate choice. Again, whether you
know it or not, you are already familiar with the important theorem
/integraldisplay
MdH=/integraldisplay
∂MH (6)
withHap-form and ∂M the boundary of a (p+1)-dimensional manifold M.
Dirac quantization of magnetic charge
After this dose of mathematics, we are ready to do some physics. Consider a sphere
surrounding a magnetic monopole with magnetic charge g. Then the electromagnetic
field 2-form is given by F=(g/4π)d cosθd ϕ . This is almost a definition of what we mean
by a magnetic monopole (see exercise IV .4.3.) In particular, calculate the magnetic flux byintegrating Fover the sphere S
2
/integraldisplay
S2F=g (7)
As I have already noted, the area element is automatically included. Indeed, you might
have recognized dcosθd ϕ=−sin θd θd ϕ as precisely the area element on a unit sphere.
Note that in “ordinary notation” (7) implies the magnetic field /vectorB=(g/4πr2)ˆr, with ˆrthe
unit vector in the radial direction.
I will now give a rather mathematical, but rigorous, derivation, originally developed by
Wu and Y ang, of Dirac’s quantization of the magnetic charge g.
First, let us recall how gauge invariance works, from, for example, (II.7.3). Under a
transformation of the electron field ψ(x)→ei/Lambda1(x)ψ(x) , the electromagnetic gauge poten-
tial changes by
Aμ(x)→Aμ(x)+1
iee−i/Lambda1(x)∂μei/Lambda1(x)
or in the language of forms,
A→A+1
iee−i/Lambda1dei/Lambda1(8)
Differentiating, we can of course write
Aμ(x)→Aμ(x)+1
e∂μ/Lambda1(x)
as is commonly done. The form given in (8) reminds us that gauge transformation is
defined as multiplication by a phase factor ei/Lambda1(x), so that /Lambda1(x) and/Lambda1(x)+2πdescribe
exactly the same transformation.
In quantum mechanics Ais physical, pace Heaviside, and so we should ask what A
would give rise to F=(g/4π)d cosθd ϕ . Easy, you say; clearly A=(g/4π)cosθd ϕ . (In
checking this by calculating dA, remember that dd=0.)
But not so fast; your mathematician friend says that dϕis not defined at the north and
south poles. Put his objection into everyday language: If you are standing on the north pole,what is your longitude? So strictly speaking it is forbidden to write A=(g/4π)cosθd ϕ .
IV .4. Magnetic Monopole | 249
But, you are smart enough to counter, then what about AN=(g/4π)(cosθ−1)dϕ , eh?
When you act with donANyou obtain the desired F; the added piece (g/4π)(−1)dϕ gets
annihilated by dthanks once again to the identity (4). At the north pole, cos θ=1,AN
vanishes, and is thus perfectly well defined.
OK, but your mathematician friend points out that your ANis not defined at the south
pole, where it is equal to (g/4π)(−2)dϕ .
Right, you respond, I anticipated that by adding the subscript N. I am now also forced
to define AS=(g/4π)(cosθ+1)dϕ . Note that dacting on ASagain gives the desired F.
But now ASis defined everywhere except at the north pole.
In mathematical jargon, we say that the gauge potential Ais defined locally, but not
globally. The gauge potential ANis defined on a “coordinate patch” covering the northern
hemisphere and extending past the equator as far south as we want as long as we donot include the south pole. Similarly, A
Sis defined on a “coordinate patch” covering the
southern hemisphere and extending past the equator as far north as we want as long aswe do not include the north pole.
But what happens where the two coordinate patches overlap, for example, along the
equator. The gauge potentials A
NandASare not the same:
AS−AN=2g
4πdϕ (9)
Now what? Aha, but this is a gauge theory: If ASandANare related by a gauge transforma-
tion, then all is well. Thus, referring to (8) we require that 2 (g/4π)dϕ =(1/ie)e−i/Lambda1dei/Lambda1
for some phase function ei/Lambda1. By inspection we have ei/Lambda1=ei2(eg/4π)ϕ.
Butϕ=0 and ϕ=2πdescribe exactly the same point. In order for ei/Lambda1to make sense,
we must have ei2(eg/4π)(2π)=ei2(eg/4π)(0)=1; in other words, eieg=1, or
g=2π
en (10)
where ndenotes an integer. This is Dirac’s famous discovery that the magnetic charge on
a magnetic monopole is quantized in units of 2 π/e . A “dual” way of putting this is that if
the monopole exists then electric charge is quantized in units of 2 π/g .
Note that the whole point is that Fis locally but not globally exact; otherwise by (6) the
magnetic charge g=/integraltext
S2Fwould be zero.
I show you this rigorous mathematical derivation partly to cut through a lot of the
confusion typical of the derivations in elementary texts and partly because this type ofargument is used repeatedly in more advanced areas of physics, such as string theory.
Electromagnetic duality
That a duality may exist between electric and magnetic fields has tantalized theoretical
physicists for a century and a half. By the way, if you read Maxwell, you will discover that heoften talked about magnetic charges. You can check that Maxwell’s equations are invariantunder the elegant transformation (/vectorE+i/vectorB)→e
iθ(/vectorE+i/vectorB)if magnetic charges exist.
250 | IV . Symmetry and Symmetry Breaking
(a) (b)xμ(σ,τ) xμ(τ)
Figure IV .4.1
One intriguing feature of (10)is that if eis small, then gis large, and vice versa. What
would magnetic charges look like if they exist? They wouldn’t look any different fromelectric charges: They too interact with a 1 /rpotential, with likes repelling and opposites
attracting. In principle, we could have perfectly formulated electromagnetism in terms ofmagnetic charges, with magnetic and electric fields exchanging their roles, but the theorywould be strongly coupled, with the coupling grather than e.
Theoretical physicists are interested in duality because it allows them a glimpse into
field theories in the strongly coupled regime. Under duality, a weakly coupled field theoryis mapped into a strongly coupled field theory. This is exactly the reason why the discoverysome years ago that certain string theories are dual to others caused such enormousexcitement in the string theory community: We get to know how string theories behave inthe strongly coupled regime. More on duality in chapter VI.3.
Forms and geometry
The geometric character of differential forms is further clarified by thinking about theelectromagnetic current of a charged particle tracing out the world line X
μ(τ) inD-
dimensional spacetime (see figure IV .4.1a):
Jμ(x)=/integraldisplay
dτdXμ
dτδ(D)[x−X(τ) ] (11)
The interpretation of this elementary formula from electromagnetism is clear: dXμ/dτ is
the 4-velocity at a given value of the parameter τ(“proper time”) and the delta function
ensures that the current at xvanishes unless the particle passes through x. Note that Jμ(x)
is invariant under the reparametrization τ→τ/prime(τ).
The generalization to an extended object is more or less obvious. Consider a string. It
traces out a world sheet Xμ(τ,σ)in spacetime (see figure 1b), where σis a parameter
IV .4. Magnetic Monopole | 251
telling us where we are along the length of the string. [For example, for a closed string,
σis conventionally taken to range between 0 and 2 πwithXμ(τ,0)=Xμ(τ,2π).] The
current associated with the string is evidently given by
Jμν(x)=/integraldisplay
dτdσ det/parenleftBigg∂τXμ∂τXν
∂σXμ∂σXν/parenrightBigg
δ(D)[x−X(τ ,σ)] (12)
where ∂τ≡∂/∂τ and so forth. The determinant is forced on us by the requirement of
invariance under reparametrization τ→τ/prime(τ,σ),σ→σ/prime(τ,σ). It follows that Jμνis an
antisymmetric tensor. Hence, the analog of the electromagnetic potential Aμcoupling to
the current Jμis an antisymmetric tensor field Bμνcoupling to the current Jμν. Thus,
string theory contains a 2-form potential B=1
2Bμνdxμdxνand the corresponding 3-form
fieldH=dB. In fact, string theory typically contains numerous p-forms.
Aharonov-Bohm effect
The reality of the gauge potential Awas brought home forcefully in 1959 by Aharonov and
Bohm. Consider a magnetic field Bconfined to a region /Omega1as illustrated in figure IV .4.2.
The quantum physics of an electron is described by solving the Schr ¨odinger equation (1).
In Feynman’s path integral formalism the amplitude associated with a path Pis modified
by a multiplicative factor eie/integraltext
P/vectorA.d/vectorx, where the line integral is evaluated along the path P.
Thus, in the path integral calculation of the probability for an electron to propagate fromatob(fig. IV .4.2), there will be interference between the contributions from path 1 and
path 2 of the form
/parenleftbigg
eie/integraltext
P1/vectorA.d/vectorx/parenrightbigg/parenleftbigg
eie/integraltext
P2/vectorA.d/vectorx/parenrightbigg∗
=/parenleftbigg
eie/contintegraltext/vectorA.d/vectorx/parenrightbigg
ab
B = 0B = 0
Ω
P2P1
Figure IV .4.2
252 | IV . Symmetry and Symmetry Breaking
but/contintegraltext/vectorA.d/vectorx=/integraltext/vectorB.d/vectorSis precisely the flux enclosed by the closed curve ( P1−P2), namely
the curve going from atobalong P1and then returning from btoaalong (−P2)since
complex conjugation in effect reverses the direction of the path P2. Remarkably, the
electron feels the effect of the magnetic field even though it never wanders into a regionwith a magnetic field present.
When the Aharonov-Bohm paper was first published, no less an authority than Niels
Bohr was deeply disturbed. The effect has since been conclusively demonstrated in a seriesof beautiful experiments by T onomura and collaborators.
Coleman once told of a gedanken prank that connects the Aharonov-Bohm effect to
Dirac quantization of magnetic charge. Let us fabricate an extremely thin solenoid so thatit is essentially invisible and thread it into the lab of an unsuspecting experimentalist,perhaps our friend from chapter III.1. We turn on a current and generate a magnetic fieldthrough the solenoid. When the experimentalist suddenly sees the magnetic flux comingout of apparently nowhere, she gets so excited that she starts planning to go to Stockholm.
What is the condition that prevents the experimentalist from discovering the prank? A
careful experimentalist might start scattering electrons around to see if she can detect asolenoid. The condition that she does not see an Aharonov-Bohm effect and thus doesnot discover the prank is precisely that the flux going through the solenoid is an integertimes 2 π/e . This implies that the apparent magnetic monopole has precisely the magnetic
charge predicted by Dirac!
Exercises
IV .4.1 Prove dd=0.
IV .4.2 Show by writing out the components explicitly that dF=0 expresses something that you are familiar
with but disguised in a compact notation.
IV .4.3 Consider F=(g/4π) d cosθd ϕ . By transforming to Cartesian coordinates show that this describes a
magnetic field pointing outward along the radial direction.
IV .4.4 Restore the factors of /planckover2piandcin Dirac’s quantization condition.
IV .4.5 Write down the reparametrization-invariant current Jμνλof a membrane.
IV .4.6 Letg(x) be the element of a group G. The 1-form v=gdg†is known as the Cartan-Maurer form.
Then tr vNis trivially closed on an N-dimensional manifold since it is already an N-form. Consider
Q=/integraltext
SNtrvNwithSNtheN-dimensional sphere. Discuss the topological meaning of Q. These con-
siderations will become important later when we discuss topology in field theory in chapter V .7. [Hint:Study the case N=3 and G=SU( 2).]
IV.5 Nonabelian Gauge Theory
Most such ideas are eventually discarded or shelved. But some
persist and may become obsessions. Occasionally an obsessiondoes finally turn out to be something good.
—C. N. Y ang talking about an idea that he first had as a
student and that he kept coming back to year after year.
1
Local transformation
It was quite a nice little idea.
T o explain the idea Y ang was talking about, recall our discussion of symmetry in
chapter I.10. For the sake of definiteness let ϕ(x)={ϕ1(x),ϕ2(x),... ,ϕN(x)} be an N-
component complex scalar field transforming as ϕ(x)→Uϕ(x), with Uan element of
SU(N) . Since ϕ†→ϕ†U†andU†U=1, we have ϕ†ϕ→ϕ†ϕand∂ϕ†∂ϕ→∂ϕ†∂ϕ. The
invariance of the Lagrangian L=∂ϕ†∂ϕ−V( ϕ†ϕ)under SU(N) is obvious for any poly-
nomial V.
In the theoretical physics community there are many more people who can answer well-
posed questions than there are people who can pose the truly important questions. Thelatter type of physicist can invariably also do much of what the former type can do, but thereverse is certainly not true.
In 1954 C.N. Y ang and R. Mills asked what will happen if the transformation varies from
place to place in spacetime, or in other words, if U=U(x) is a function of x.
Clearly, ϕ
†ϕis still invariant. But in contrast ∂ϕ†∂ϕis no longer invariant. Indeed,
∂μϕ→∂μ(Uϕ)=U∂μϕ+(∂μU)ϕ=U[∂μϕ+(U†∂μU)ϕ ]
T o cancel the unwanted term (U†∂μU)ϕ , we generalize the ordinary derivative ∂μto a
covariant derivative Dμ, which when acting on ϕ, gives
Dμϕ(x)=∂μϕ(x)−iAμ(x)ϕ(x) (1)
The field Aμis called a gauge potential in direct analogy with electromagnetism.
1C. N. Y ang, Selected Papers 1945–1980 with Commentary ,p .1 9 .
254 | IV . Symmetry and Symmetry Breaking
How must Aμtransform, so that Dμϕ(x)→U(x)D μϕ(x) ? In other words, we would
likeDμϕ(x) to transform the way ∂μϕ(x) transformed when Udid not depend on x.I f
so, then [ Dμϕ(x) ]†Dμϕ(x)→[Dμϕ(x) ]†Dμϕ(x) and can be used as an invariant kinetic
energy term for the field ϕ.
Working backward, we see that Dμϕ(x)→U(x)D μϕ(x) if ( and it goes without saying
that you should be checking this)
Aμ→UAμU†−i(∂μU)U†=UAμU†+iU∂μU†(2)
(The equality follows from UU†=1.)We refer to Aμas the nonabelian gauge potential
and to (2) as a nonabelian gauge transformation.
Let us now make a series of simple observations.
1. Clearly, Aμhave to be NbyNmatrices. Work out the transformation law for A†
μusing
(2) and show that the condition Aμ−A†
μ=0 is preserved by the gauge transformation.
Thus, it is consistent to take Aμto be hermitean. Specifically, you should work out what
this means for the group SU( 2)so that U=eiθ.τ/2where θ.τ=θaτa, with τathe familiar
Pauli matrices.
2. Writing U=eiθ.TwithTathe generators of SU(N), we have
Aμ→Aμ+iθa[Ta,Aμ]+∂μθaTa(3)
under an infinitesimal transformation U/similarequal1+iθ.T. For most purposes, the infinitesimal
form (3) suffices.
3. T aking the trace of (3) we see that the trace of Aμdoes not transform and so we can take Aμ
to be traceless as well as hermitean. This means that we can always write Aμ=Aa
μTaand
thus decompose the matrix field Aμinto component fields Aa
μ. There are as many Aa
μ’s as
there are generators in the group [3 for SU( 2), 8 for SU( 3), and so forth.]
4. You are reminded in appendix B that the Lie algebra of the group is defined by [ Ta,Tb]=
ifabcTc, where the numbers fabcare called structure constants. For example, fabc=εabc
forSU( 2). Thus, (3) can be written as
Aa
μ→Aa
μ−fabcθbAc
μ+∂μθa(4)
Note that if θdoes not depend on x, theAa
μ’s transform as the adjoint representation of the
group.
5. IfU(x)=eiθ(x)is just an element of the abelian group U(1), all these expressions simplify
andAμis just the abelian gauge potential familiar from electromagnetism, with (2 ) the usual
abelian gauge transformation. Hence, Aμis known as the nonabelian gauge potential.
A transformation Uthat depends on the spacetime coordinates xis known as a gauge
transformation or local transformation. A Lagrangian Linvariant under a gauge transfor-
mation is said to be gauge invariant.
IV .5. Nonabelian Gauge Theory | 255
Construction of the field strength
We can now immediately write a gauge invariant Lagrangian, namely
L=(D μϕ)†(Dμϕ)−V( ϕ†ϕ) (5)
but the gauge potential Aμdoes not yet have dynamics of its own. In the familiar example
ofU(1)gauge invariance, we have written the coupling of the electromagnetic potential
Aμto the matter field ϕ, but we have yet to write the Maxwell term −1
4FμνFμνin the
Lagrangian. Our first task is to construct a field strength Fμνout of Aμ. How do we do
that? Y ang and Mills apparently did it by trial and error. As an exercise you might alsowant to try that before reading on.
At this point the language of differential forms introduced in chapter IV .4 proves to
be of use. It is convenient to absorb a factor of −i by defining A
M
μ≡−iAP
μ, where AP
μ
denotes the gauge potential we have been using all along. Until further notice, when
we write Aμwe mean AM
μ. Referring to (1) we see that the covariant derivative has the
cleaner form Dμ=∂μ+Aμ. (Incidentally, the superscripts MandPindicate the potential
appearing in the mathematical and physical literature, respectively.) As before, let usintroduce A=A
μdxμ, now a matrix 1-form, that is, a form that also happens to be a matrix
in the defining representation of the Lie algebra [e.g., an NbyNtraceless hermitean matrix
forSU(N).] Note that
A2=AμAνdxμdxν=1
2[Aμ,Aν]dxμdxν
is not zero for a nonabelian gauge potential. (Obviously, there is no such object in electro-
magnetism.)
Our task is to construct a 2-form F=1
2Fμνdxμdxνout of the 1-form A. We adopt a direct
approach. Out of Awe can construct only two possible 2-forms: dA andA2.S oFmust be
a linear combination of the two.
In the notation we are using the transformation law (2) reads
A→UAU†+UdU†(6)
withUa 0-form (and so dU†=∂μU†dxμ.)Applying dto (6) we have
dA→UdAU†+dUAU†−UAdU†+dUdU†(7)
Note the minus sign in the third term, from moving the 1-form dpast the 1-form A.O n
the other hand, squaring (6) we have
A2→UA2U†+UAdU†+UdU†UAU†+UdU†UdU†(8)
Applying dtoUU†=1 we have UdU†=−dUU†. Thus, we can rewrite (8)as
A2→UA2U†+UAdU†−dUAU†−dUdU†(9)
256 | IV . Symmetry and Symmetry Breaking
Lo and behold! If we add (7) and (9), six terms knock each other off, leaving us with
something nice and clean:
dA+A2→U(dA +A2)U†(10)
The mathematical structure thus led Y ang and Mills to define the field strength
F=dA+A2(11)
Unlike A, the field strength 2-form Ftransforms homogeneously (10):
F→UFU†(12)
In the abelian case A2vanishes and Freduces to the usual electromagnetic form. In the
nonabelian case, Fis not gauge invariant, but gauge covariant.
Of course, you can also construct Fa
μνwithout using differential forms. As an exercise
you should do it starting with (4). The exercise will make you appreciate differential forms!
At the very least, we can regard differential forms as an elegantly compact notation thatsuppresses the indices aandμin (4). At the same time, the fact that (11) emerges so
smoothly clearly indicates a profound underlying mathematical structure. Indeed, thereis a one-to-one translation between the physicist’s language of gauge theory and themathematician’s language of fiber bundles.
Let me show you another route to (11). In analogy to d, define D=d+A, understood
as an operator acting on a form to its right. Let us calculate
D2=(d+A)(d+A)=d2+dA+Ad+A2
The first term vanishes, the second can be written as dA=(dA)−Ad; the parenthesis
emphasizes that dacts only on A. Thus,
D2=(dA)+A2=F (13)
Pretty slick? I leave it as an exercise for you to show that D2transforms homogeneously
and hence so does F.
Elegant though differential forms are, in physics it is often desirable to write more
explicit formulas. We can write (11) out as
F=(∂μAν+AμAν)dxμdxν=1
2(∂μAν−∂νAμ+[Aμ,Aν])dxμdxν(14)
With the definition F≡1
2Fμνdxμdxνwe have
Fμν=∂μAν−∂νAμ+[Aμ,Aν] (15)
At this point, we might also want to switch back to physicist’s notation. Recall that Aμ
in (15) is actually AM
μ≡−iAP
μand so by analogy define FM
μν=−iFP
μν. Thus,
Fμν=∂μAν−∂νAμ−i[Aμ,Aν] (16)
where, until further notice, Aμstands for AP
μ. (One way to see the necessity for the iin (16)
is to remember that physicists like to take Aμto be a hermitean matrix and the commutator
of two hermitean matrices is antihermitean.)
IV .5. Nonabelian Gauge Theory | 257
As long as we are being explicit we might as well go all the way and exhibit the group
indices as well as the Lorentz indices. We already wrote Aμ=Aa
μTaand so we naturally
writeFμν=Fa
μνTa. Then (16) becomes
Fa
μν=∂μAa
ν−∂νAa
μ+fabcAb
μAcν (17)
I mention in passing that for SU( 2)Aand Ftransform as vectors and the structure
constant fabcis just εabc, so the vector notation /vectorFμν=∂μ/vectorAν−∂ν/vectorAμ+/vectorAμ×/vectorAνis often
used.
The Y ang-Mills Lagrangian
Given that Ftransforms homogeneously (12) we can immediately write down the analog
of the Maxwell Lagrangian, namely the Y ang-Mills Lagrangian
L=−1
2g2trFμνFμν(18)
We are normalizing Taby trTaTb=1
2δabso that L=−(1/4g2)Fa
μνFaμν. The theory
described by this Lagrangian is known as pure Y ang-Mills theory or nonabelian gaugetheory.
Apart from the quadratic term (∂
μAa
ν−∂νAa
μ)2, the Lagrangian L=−(1/4g2)Fa
μνFaμν
also contains a cubic term fabcAbμAcν(∂μAa
ν−∂νAa
μ)and a quartic term (fabcAb
μAcν)2.
As in electromagnetism the quadratic term describes the propagation of a massless vector
boson carrying an internal index a, known as the nonabelian gauge boson or the Y ang-
Mills boson. The cubic and quartic terms are not present in electromagnetism and describethe self-interaction of the nonabelian gauge boson. The corresponding Feynman rules aregiven in figure IV .5.1a, 1b, and c.
The physics behind this self-interaction of the Y ang-Mills bosons is not hard to under-
stand. The photon couples to charged fields but is not charged itself. Just as the charge
(c) (b)(a)
Figure IV .5.1
258 | IV . Symmetry and Symmetry Breaking
of a field tells us how the field transforms under the U(1)gauge group, the analog of the
charge of a field in a nonabelian gauge theory is the representation the field belongs to. TheY ang-Mills bosons couple to all fields transforming nontrivially under the gauge group. Butthe Y ang-Mills bosons themselves transform nontrivially: In fact, as we have noted, theytransform under the adjoint representation. Thus, they must couple to themselves.
Pure Maxwell theory is free and so essentially trivial. It contains a noninteracting photon.
In contrast, pure Y ang-Mills theory contains self-interaction and is highly nontrivial. Notethat the structure coefficients f
abcare completely fixed by group theory, and thus in
contrast to a scalar field theory, the cubic and quartic self-interactions of the gauge bosons,including their relative strengths, are totally fixed by symmetry. If any 4-dimensional fieldtheory can be solved exactly, pure Y ang-Mills theory may be it, but in spite of the enormousamount of theoretical work devoted to it, it remains unsolved (see chapters VII.3 and VII.4).
’t Hooft’s double-line formalism
While it is convenient to use the component fields Aa
μfor many purposes, the matrix
fieldAμ=Aa
μTaembodies the mathematical structure of nonabelian gauge theory more
elegantly. The propagator for the components of the matrix field in a U(N) gauge theory
has the form
/angbracketleft0|TAμ(x)i
jAν(0)k
l|0/angbracketright
=/angbracketleft0|TAa
μ(x)Abν(0)|0/angbracketright(Ta)i
j(Tb)kl (19)
∝δab(Ta)i
j(Tb)kl∝δi
lδk
j
[We have gone from an SU(N) to aU(N) theory for the sake of simplicity. The generators
ofSU(N) satisfy a traceless condition T r Ta=0, as a result of which we would have to
subtract1
Nδi
jδk
lfrom the right-hand side.] The matrix structure Ai
μjnaturally suggests that
we, following ’t Hooft, introduce a double-line formalism, in which the gauge potential isdescribed by two lines, each associated with one of the two indices iandj. We choose the
convention that the upper index flows into the diagram, while the lower index flows outof the diagram. The propagator in (19) is represented in figure IV .5.2a. The double-lineformalism allows us to reproduce the index structure δ
i
lδk
jnaturally. The cubic and quartic
couplings are represented in figure IV .5.2b and c.
The constant gintroduced in (18) is known as the Y ang-Mills coupling constant. We can
always write the quadratic term in (18) in the convention commonly used in electromag-netism by a trivial rescaling A→gA. After this rescaling, the cubic and quartic couplings
of the Y ang-Mills boson go as gandg
2, respectively. The covariant derivative in (1) be-
comes Dμϕ=∂μϕ−igAμϕ, showing that galso measures the coupling of the Y ang-Mills
boson to matter. The convention we used, however, brings out the mathematical structure
more clearly. As written in (18), g2measures the ease with which the Y ang-Mills boson can
propagate. Recall that in chapter III.7 we also found this way of defining the coupling as
IV .5. Nonabelian Gauge Theory | 259
i
jl
k
(a)
(b) (c)
Figure IV .5.2
a measure of propagation useful in electromagnetism. We will see in chapter VIII.1 that
Newton’s coupling appears in the same way in the Einstein-Hilbert action for gravity.
Theθterm
Besides tr FμνFμν, we can also form the dimension-4 term εμνλρtrFμνFλρ. Clearly,
this term violates time reversal invariance Tand parity Psince it involves one time
index and three space indices. We will see later that the strong interaction is describedby a nonabelian gauge theory, with the Lagrangian containing the so-called θterm
(θ/32π
2)εμνλρtrFμνFλρ. As you will show in exercise IV .5.3 this term is a total diver-
gence and does not contribute to the equation of motion. Nevertheless, it induces anelectric dipole moment for the neutron. The experimental upper bound on the electricdipole moment for the neutron translates into an upper bound on θof the order 10
−9.I
will not go into how particle physicists resolve the problem of making sure that θis small
enough or vanishes outright.
Coupling to matter fields
We took the scalar field ϕto transform in the fundamental representation of the group.
In general, ϕcan transform in an arbitrary representation Rof the gauge group G.W e
merely have to write the covariant derivative more generally as
Dμϕ=(∂μ−iAa
μTa
(R))ϕ (20)
where Ta
(R)represents the ath generator in the representation R(see exercise IV .5.1).
Clearly, the prescription to turn a globally symmetric theory into a locally symmetric
theory is to replace the ordinary derivative ∂μacting on any field, boson or fermion,
260 | IV . Symmetry and Symmetry Breaking
belonging to the representation Rby the covariant derivative Dμ=(∂μ−iAa
μTa
(R)). Thus,
the coupling of the nonabelian gauge potential to a fermion field is given by
L=¯ψ(iγμDμ−m)ψ=¯ψ(iγμ∂μ+γμAa
μTa
(R)−m)ψ . (21)
Fields listen to the Y ang-Mills gauge bosons according to the representation Rthat
they belong to, and those that belong to the trivial identity representation do not hearthe call of the gauge bosons. In the special case of a U(1)gauge theory, also known as
electromagnetism, Rcorresponds to the electric charge of the field. Those fields that
transform trivially under U(1)are electrically neutral.
Appendix
Let me show you another context, somewhat surprising at first sight, in which the Y ang-Mills structure pops up.2
Consider Schr ¨odinger’s equation
i∂
∂t/Psi1(t)=H(t)/Psi1(t) (22)
with a time dependent Hamiltonian H(t) . The setup is completely general: For instance, we could be talking
about spin states in a magnetic field or about a single particle nonrelativistic Hamiltonian with the wave function/Psi1(/vectorx,t). We suppress the dependence of Hand/Psi1on variables other than time t.
First, solve the eigenvalue problem of H(t) . Suppose that because of symmetry or some other reason the
spectrum of H(t ) contains an n-fold degeneracy, in other words, there exist ndistinct solutions of the equations
H(t)ψ
a(t)=E(t)ψa(t), with a=1,... ,n. Note that E(t) can vary with time and that we are assuming that the
degeneracy persists with time, that is, the degeneracy does not occur “accidentally” at one instant in time. We canalways replace H(t) byH(t)−E(t) so that henceforth we have H(t)ψ
a(t)=0. Also, the states can be chosen to
be orthogonal so that /angbracketleftψb(t)|ψ a(t)/angbracketright=δ ba. (For notational reasons it is convenient to jump back and forth between
the Schr ¨odinger and the Dirac notation. T o make it absolutely clear, we have
/angbracketleftψb(t)|ψ a(t)/angbracketright=/integraldisplay
d/vectorxψ∗
b(/vectorx,t)ψa(/vectorx,t)
if we are talking about single particle quantum mechanics.)
Let us now study (22) in the adiabatic limit, that is, we assume that the time scale over which H(t) varies is
much longer than 1 //Delta1E , where /Delta1E denotes the energy gap separating the states ψa(t)from neighboring states.
In that case, if /Psi1(t) starts out in the subspace spanned by {ψa(t)} it will stay in that subspace and we can write
/Psi1(t)=/summationtext
aca(t)ψa(t). Plugging this into (22) we obtain immediately/summationtext
a[(dca/dt)ψ a(t)+ca(t)(∂ψa/∂t)]=0.
T aking the scalar product with ψb(t), we obtain
dcb
dt=−/summationdisplay
aAbaca (23)
with the nbynmatrix
Aba(t)≡i/angbracketleftψb(t)|∂ψa
∂t/angbracketright (24)
Now suppose somebody else decides to use a different basis, ψ/prime
a(t)=U∗
ac(t)ψc(t), related to ours by a unitary
transformation. (The complex conjugate on the unitary matrix Uis just a notational choice so that our final
equation will come out looking the same as a celebrated equation in the text; see below.) I have also passed
2F. Wilczek and A. Zee, “Appearance of Gauge Structure in Simple Dynamical Systems,” Phys. Rev. Lett.
52:2111, 1984.
IV .5. Nonabelian Gauge Theory | 261
to the repeated indices summed notation. Differentiate to obtain (∂ψ/prime
a/∂t)=U∗
ac(t)(∂ψ c/∂t)+(dU∗
ac/dt)ψc(t).
Contracting this with ψ/prime∗
b(t)=Ubd(t)ψ∗
d(t)and multiplying by i,w ef i n d
A/prime=UAU†+iU∂U†
∂t(25)
Suppose the Hamiltonian H(t) depends on dparameters λ1,... ,λd. We vary the parameters, thus tracing out a
path defined by {λμ(t),μ=1,... ,d}in the d-dimensional parameter space. For example, for a spin Hamiltonian,
{λμ}could represent an external magnetic field. Now (23) becomes
dcb
dt=−/summationdisplay
a(Aμ)bacadλμ
dt(26)
if we define (Aμ)ba≡i/angbracketleftψb|∂μψa/angbracketright, where ∂μ≡∂/∂λμ, and (25) generalizes to
A/prime
μ=UAμU†+iU∂μU†(27)
We have recovered (IV .5.2). Lo and behold, a Y ang-Mills gauge potential Aμhas popped up in front of our very
eyes!
The “transport” equation (26) can be formally solved by writing c(λ)=Pe−/integraltext
Aμdλμ, where the line integral
is over a path connecting an initial point in the parameter space to some final point λandPdenotes a path
ordering operation. We break the path into infinitesimal segments and multiply together the noncommuting
contribution e−Aμ/Delta1λμfrom each segment, ordered along the path. In particular, if the path is a closed curve, by
the time we return to the initial values of the parameters, the wave function will have acquired a matrix phasefactor, known as the nonabelian Berry’s phase. This discussion is clearly intimately related to the discussion ofthe Aharonov-Bohm phase in the preceding chapter.
T o see this nonabelian phase, all we have to do is to find some quantum system with degeneracy in its spectrum
and vary some external parameter such as a magnetic field.
3In their paper, Y ang and Mills spoke of the degeneracy
of the proton and neutron under isospin in an idealized world and imagined transporting a proton from one pointin the universe to another. That a proton at one point can be interpreted as a neutron at another necessitates theintroduction of a nonabelian gauge potential. I find it amusing that this imagined transport can now be realizedanalogously in the laboratory.
You will realize that the discussion here parallels the discussion in the text leading up to (IV .5.2). The spacetime
dependent symmetry transformation corresponds to a parameter dependent change of basis. When I discussgravity in chapter VIII.1 it will become clear that moving the basis {ψ
a}around in the parameter space is the
precise analog of parallel transporting a local coordinate frame in differential geometry and general relativity. We
will also encounter the quantity Pe−/integraltext
Aμdλμagain in chapter VII.1 in the guise of a Wilson loop.
Exercises
IV .5.1 Write down the Lagrangian of an SU( 2)gauge theory with a scalar field in the I=2 representation.
IV .5.2 Prove the Bianchi identity DF≡dF+[A,F]=0. Write this out explicitly with indices and show that in
the abelian case it reduces to half of Maxwell’s equations.
IV .5.3 In 4-dimensions εμνλρtrFμνFλρcan be written as tr F2. Show that dtrF2=0 in any dimensions.
IV .5.4 Invoking the Poincar ´e lemma (IV .4.5) and the result of exercise IV .5.3 show that tr F2=dtr(AdA +2
3A3).
Write this last equation out explicitly with indices. Identify these quantities in the case of electromag-netism.
3A. Zee, “On the Non-Abelian Gauge Structure in Nuclear Quadrupole Resonance,” Phys. Rev. A38:1, 1988.
The proposed experiment was later done by A. Pines.
262 | IV . Symmetry and Symmetry Breaking
IV .5.5 For a challenge show that tr Fn, which appears in higher dimensional theories such as string theory, are
all total divergences. In other words, there exists a (2n−1)-form ω2n−1(A) such that tr Fn=dω 2n−1(A) .
[Hint: A compact representation of the form ω2n−1(A)=/integraltext1
0dt f 2n−1(t,A) exists.] Work out ω5(A)
explicitly and try to generalize knowing ω3andω5. Determine the (2n−1)-form f2n−1(t,A). For help,
see B. Zumino et al., Nucl. Phys. B239:477, 1984.
IV .5.6 Write down the Lagrangian of an SU( 3)gauge theory with a fermion field in the fundamental or defining
triplet representation.
IV.6 The Anderson-Higgs Mechanism
The gauge potential eats the Nambu-Goldstone boson
As I noted earlier the ability to ask good questions is of crucial importance in physics.
Here is an excellent question: How does spontaneous symmetry breaking manifest itselfin gauge theories?
Going back to chapter IV .1, we gauge the U(1)theory in (IV .1.6) by replacing ∂
μϕwith
Dμϕ=(∂μ−ieAμ)ϕso that
L=−1
4FμνFμν+(Dϕ)†Dϕ+μ2ϕ†ϕ−λ(ϕ†ϕ)2(1)
Now when we go to polar coordinates ϕ=ρeiθwe have Dμϕ=[∂μρ+iρ(∂μθ−eAμ)]eiθ
and thus
L=−1
4FμνFμν+ρ2(∂μθ−eAμ)2+(∂ρ)2+μ2ρ2−λρ4(2)
(Compare this with L=ρ2(∂μθ)2+(∂ρ)2+μ2ρ2−λρ4in the absence of the gauge field.)
Under a gauge transformation ϕ→eiαϕ(so that θ→θ+α)andeAμ→eAμ+∂μα, and
thus the combination Bμ≡Aμ−(1/e)∂μθis gauge invariant. The first two terms in L
thus become −1
4FμνFμν+e2ρ2B2
μ. Note that Fμν=∂μAν−∂νAμ=∂μBν−∂νBμhas the
same form in terms of the potential Bμ.
Upon spontaneous symmetry breaking, we write ρ=(1/√
2)(v+χ), with v=/radicalbig
μ2/λ.
Hence
L=−1
4FμνFμν+1
2M2B2
μ+e2vχB2
μ+1
2e2χ2B2
μ
+1
2(∂χ)2−μ2χ2−√
λμχ3−λ
4χ4+μ4
4λ(3)
264 | IV . Symmetry and Symmetry Breaking
The theory now consists of a vector field Bμwith mass
M=ev (4)
interacting with a scalar field χwith mass√
2μ. The phase field θ, which would have been
the Nambu-Goldstone boson in the ungauged theory, has disappeared. We say that thegauge field A
μhas eaten the Nambu-Goldstone boson; it has gained weight and changed
its name to Bμ.
Recall that a massless gauge field has only 2 degrees of freedom, while a massive gauge
field has 3 degrees of freedom. A massless gauge field has to eat a Nambu-Goldstone bosonin order to have the requisite number of degrees of freedom. The Nambu-Goldstoneboson becomes the longitudinal degree of freedom of the massive gauge field. We do notlose any degrees of freedom, as we had better not.
This phenomenon of a massless gauge field becoming massive by eating a Nambu-
Goldstone boson was discovered by numerous particle physicists
1and is known as the
Higgs mechanism. People variously call ϕ, or more restrictively χ, the Higgs field. The
same phenomenon was discovered in the context of condensed matter physics by Landau,Ginzburg, and Anderson, and is known as the Anderson mechanism.
Let us give a slightly more involved example, an O(3)gauge theory with a Higgs field ϕ
a
(a=1, 2, 3)transforming in the vector representation. The Lagrangian contains the kinetic
energy term1
2(Dμϕa)2, with Dμϕa=∂μϕa+gεabcAb
μϕcas indicated in (IV .5.20). Upon
spontaneous symmetry breaking, /vectorϕacquires a vacuum expectation value which without
loss of generality we can choose to point in the 3-direction, so that /angbracketleftϕa/angbracketright=vδa3. We set
ϕ3=vand see that
1
2(Dμϕa)2→1
2(gv)2(A1
μAμ1+A2
μAμ2) (5)
The gauge potential A1
μandA2
μacquires mass gv[compare with (4)] while A3
μremains
massless.
A more elaborate example is that of an SU( 5)gauge theory with ϕtransforming as the 24-
dimensional adjoint representation. (See appendix B for the necessary group theory.) Thefieldϕi sa5b y5hermitean traceless matrix. Since the adjoint representation transforms
asϕ→ϕ+iθ
a[Ta,ϕ], we have Dμϕ=∂μϕ−igAa
μ[Ta,ϕ] with a=1,... , 24 running
over the 24 generators of SU( 5). By a symmetry transformation the vacuum expectation
value of ϕcan be taken to be diagonal /angbracketleftϕi
j/angbracketright=vjδi
j(i,j=1,... ,5), with/summationtext
jvj=0. (This
is the analog of our choosing /angbracketleft/vectorϕ/angbracketrightto point in the 3-direction in the preceding example.) We
have in the Lagrangian
tr(Dμϕ)(Dμϕ)→g2tr[Ta,/angbracketleftϕ/angbracketright][/angbracketleftϕ/angbracketright,Tb]Aa
μAμb(6)
The gauge boson masses squared are given by the eigenvalues of the 24 by 24 matrix
g2tr[Ta,/angbracketleftϕ/angbracketright][/angbracketleftϕ/angbracketright,Tb], which we can compute laboriously for any given /angbracketleftϕ/angbracketright.
1Including P. Higgs, F. Englert, R. Brout, G. Guralnik, C. Hagen, and T . Kibble.
IV .6. Anderson-Higgs Mechanism | 265
It is easy to see, however, which gauge bosons remain massless. As a specific example
(which will be of interest to us in chapter VII.6), suppose
/angbracketleftϕ/angbracketright=v⎛
⎜⎜⎜⎜⎜⎜⎜⎜⎝200 0 0
020 0 0002 0 0000−30000 0 −3⎞
⎟⎟⎟⎟⎟⎟⎟⎟⎠(7)
Which generators Tacommute with /angbracketleftϕ/angbracketright? Clearly, generators of the form/parenleftBig
A0
00/parenrightBig
and of the
form/parenleftBig
00
0B/parenrightBig
. Here Arepresents 3 by 3 hermitean traceless matrices (of which there are
32−1=8, the so-called Gell-Mann matrices) and Brepresents 2 by 2 hermitean traceless
matrices (of which there are 22−1=3, namely the Pauli matrices). Furthermore, the
generator
⎛
⎜⎜⎜⎜⎜⎜⎜⎜⎝200 0 0
020 0 0002 0 0000−30000 0 −3⎞
⎟⎟⎟⎟⎟⎟⎟⎟⎠(8)
being proportional to /angbracketleftϕ/angbracketright, obviously commutes with /angbracketleftϕ/angbracketright. Clearly, these generators gen-
erateSU( 3),SU( 2), and U(1), respectively. Thus, in the 24 by 24 mass-squared matrix
g2tr[Ta,/angbracketleftϕ/angbracketright][/angbracketleftϕ/angbracketright,Tb] there are blocks of submatrices that vanish, namely, an 8 by 8 block,
a 3 by 3 block, anda1b y1block. We have 8 +3+1=12 massless gauge bosons. The
remaining 24 −12=12 gauge bosons acquire mass.
Counting massless gauge bosons
In general, consider a theory with the global symmetry group Gspontaneously broken to
a subgroup H. As we learned in chapter IV .1, n(G)−n(H) Nambu-Goldstone bosons
appear. Now suppose the symmetry group Gis gauged. We start with n(G) massless
gauge bosons, one for each generator. Upon spontaneous symmetry breaking, the n(G)−
n(H) Nambu-Goldstone bosons are eaten by n(G)−n(H) gauge bosons, leaving n(H)
massless gauge bosons, exactly the right number since the gauge bosons associated withthe surviving gauge group Hshould remain massless.
In our simple example, G=U(1),H=nothing: n(G)=1 andn(H)=0. In our second
example, G=O(3),H=O(2)/similarequalU(1):n(G)=3 and n(H)=1, and so we end up with
one massless gauge boson. In the third example, G=SU( 5),H=SU( 3)⊗SU( 2)⊗U(1)
so that n(G)=24 and n(H)=12. Further examples and generalizations are worked out
in the exercises.
266 | IV . Symmetry and Symmetry Breaking
Gauge boson mass spectrum
It is easy enough to work out the mass spectrum explicitly. The covariant derivative of a
Higgs field is Dμϕ=∂μϕ+gAa
μTaϕ, where gis the gauge coupling, Taare the generators
of the group Gwhen acting on ϕ, andAa
μthe gauge potential corresponding to the ath
generator. Upon spontaneous symmetry breaking we replace ϕby its vacuum expectation
value/angbracketleftϕ/angbracketright=v. Hence Dμϕis replaced by gAa
μTav. The kinetic term1
2(Dμϕ.Dμϕ)[here
(.)denotes the scalar product in the group G] in the Lagrangian thus becomes
1
2g2(Tav.Tbv)AμaAb
μ≡1
2Aμa(μ2)abAbμ
where we have introduced the mass-squared matrix
(μ2)ab=g2(Tav.Tbv) (9)
for the gauge bosons. [You will recognize (9) as the generalization of (4); also compare (5)
and (6).] We diagonalize (μ2)abto obtain the masses of the gauge bosons. The eigenvectors
tell us which linear combinations of Aa
μcorrespond to mass eigenstates.
Note that μ2is ann(G) byn(G) matrix with n(H) zero eigenvalues, whose existence can
also be seen explicitly. Let Tcbe a generator of H. The statement that Hremains unbroken
by the vacuum expectation value vmeans that the symmetry transformation generated by
Tcleaves vinvariant; in other words, Tcv=0, and hence the gauge boson associated with
Tcremains massless, as it should. All these points are particularly evident in the SU( 5)
example we worked out.
Feynman rules in spontaneously broken gauge theories
It is easy enough to derive the Feynman rules for spontaneously broken gauge theories.T ake, for example, (3). As usual, we look at the terms quadratic in the fields, Fouriertransform, and invert. We see that the gauge boson propagator is given by
−i
k2−M2+iε(gμν−kμkν
M2) (10)
and the χpropagator by
i
k2−2μ2+iε(11)
I leave it to you to work out the rules for the interaction vertices.
As I said in another context, field theories often exist in several equivalent forms.
T ake the U(1)theory in (1) and instead of polar coordinates go to Cartesian coordinates
ϕ=(1/√
2)(ϕ1+iϕ2)so that
Dμϕ=∂μϕ−ieAμϕ=1√
2[(∂μϕ1+eAμϕ2)+i(∂μϕ2−eAμϕ1)]
IV .6. Anderson-Higgs Mechanism | 267
Then (1) becomes
L=−1
4FμνFμν+1
2[(∂μϕ1+eAμϕ2)2+(∂μϕ2−eAμϕ1)2] (12)
+1
2μ2(ϕ2
1+ϕ2
2)−1
4λ(ϕ2
1+ϕ2
2)2
Spontaneous symmetry breaking means setting ϕ1→v+ϕ/prime
1withv=/radicalbig
μ2/λ.
The physical content of (12) and (3) should be the same. Indeed, expand the Lagrangian
(12) to quadratic order in the fields:
L=μ4
4λ−1
4FμνFμν+1
2M2A2
μ−MAμ∂μϕ2+1
2[(∂μϕ/prime
1)2−2μ2ϕ/prime2
1]
+1
2(∂μϕ2)2+... (13)
The spectrum, a gauge boson Awith mass M=evand a scalar boson ϕ/prime
1with mass√
2μ,
is identical to the spectrum in (3). (The particles there were named Bandχ.)
But oops, you may have noticed something strange: the term −MAμ∂μϕ2which mixes
the fields Aμandϕ2. Besides, why is ϕ2still hanging around? Isn’t he supposed to have
been eaten? What to do?
We can of course diagonalize but it is more convenient to get rid of this mixing term.
Referring to the Fadeev-Popov quantization of gauge theories discussed in chapter III.4 wenote that the gauge fixing term generates a term to be added to L. We can cancel the unde-
sirable mixing term by choosing the gauge function to be f (A)=∂A+ξevϕ
2−σ. Going
through the steps, we obtain the effective Lagrangian Leff=L−(1/2ξ)(∂A +ξMϕ 2)2
[compare with (III.4.7)]. The undesirable cross term −MAμ∂μϕ2inLis now canceled upon
integration by parts. In Leffthe terms quadratic in Anow read −1
4FμνFμν+1
2M2A2
μ−
(1/2ξ)(∂A)2while the terms quadratic in ϕ2read1
2[(∂μϕ2)2−ξM2ϕ22], immediately giving
us the gauge boson propagator
−i
k2−M2+iε/bracketleftbigg
gμν−(1−ξ)kμkν
k2−ξM2+iε/bracketrightbigg
(14)
and the ϕ2propagator
i
k2−ξM2+iε(15)
This one-parameter class of gauge choices is known as the Rξgauge. Note that the would-
be Goldstone field ϕ2remains in the Lagrangian, but the very fact that its mass depends on
the gauge parameter ξbrands it as unphysical. In any physical process, the ξdependence in
theϕ2andApropagators must cancel out so as to leave physical amplitudes ξindependent.
In exercise IV .6.9 you will verify that this is indeed the case in a simple example.
Different gauges have different advantages
You might wonder why we would bother with the Rξgauge. Why not just use the equivalent
formulation of the theory in (3), known as the unitary gauge, in which the gauge boson
268 | IV . Symmetry and Symmetry Breaking
propagator (10) looks much simpler than (14) and in which we don’t have to deal with the
unphysical ϕ2field? The reason is that the Rξgauge and the unitary gauge complement
each other. In the Rξgauge, the gauge boson propagator (14) goes as 1 /k2for large kand
so renormalizability can be proved rather easily. On the other hand, in the unitary gaugeall fields are physical (hence the name “unitary”) but the gauge boson propagator (10)apparently goes as k
μkν/k2for large k; to prove renormalizability we must show that the
kμkνpiece of the propagator does not contribute. Using both gauges, we can easily prove
that the theory is both renormalizable and unitary. By the way, note that in the limit ξ→∞
(14) goes over to (10) and ϕ2disappear, at least formally.
In practical calculations, there are typically many diagrams to evaluate. In the Rξgauge,
the parameter ξdarn well better disappears when we add everything up to form the physical
mass shell amplitude. The Rξgauge is attractive precisely because this requirement
provides a powerful check on practical calculations.
I remarked earlier that strictly speaking, gauge invariance is not so much a symmetry
as the reflection of a redundancy in the degrees of freedom used. (The photon has only 2degrees of freedom but we use a field A
μwith 4 components.) A purist would insist, in
the same vein, that there is no such thing as spontaneously breaking a gauge symmetry.T o understand this remark, note that spontaneous breaking amounts to setting ρ≡|ϕ|to
vandθto 0 in (2). The statement |ϕ|=v is perfectly U(1)invariant: It defines a circle in ϕ
space. By picking out the point θ=0 on the circle in a globally symmetric theory we break
the symmetry. In contrast, in a gauge theory, we can use the gauge freedom to fix θ=0
everywhere in spacetime. Hence the purists. I will refrain from such hair-splitting in thisbook and continue to use the convenient language of symmetry breaking even in a gaugetheory.
Exercises
IV .6.1 Consider an SU( 5)gauge theory with a Higgs field ϕtransforming as the 5-dimensional representation:
ϕi,i=1, 2, ... , 5. Show that a vacuum expectation value of ϕbreaks SU( 5)toSU( 4). Now add another
Higgs field ϕ/prime, also transforming as the 5-dimensional representation. Show that the symmetry can either
remain at SU( 4)or be broken to SU( 3).
IV .6.2 In general, there may be several Higgs fields belonging to various representations labeled by α. Show that
the mass squared matrix for the gauge bosons generalize immediately to (μ2)ab=/summationtext
αg2(Ta
αvα.Tb
αvα),
where vαis the vacuum expectation value of ϕαandTa
αis theath generator represented on ϕα. Combine
the situations described in exercises IV .6.1 and IV .6.2 and work out the mass spectrum of the gaugebosons.
IV .6.3 The gauge group Gdoes not have to be simple; it could be of the form G
1⊗G2⊗...⊗Gk, with
coupling constants g1,g2,... ,gk. Consider, for example, the case G=SU( 2)⊗U(1)and a Higgs
fieldϕtransforming like the doublet under SU( 2)and like a field with charge1
2under U(1), so that
Dμϕ=∂μϕ−i[gAa
μ(τa/2)+g/primeBμ1
2]ϕ. Let/angbracketleftϕ/angbracketright=/parenleftBig
0
v/parenrightBig
. Determine which linear combinations of the
gauge bosons Aa
μandBμacquire mass.
IV .6.4 In chapter IV .5 you worked out an SU( 2)gauge theory with a scalar field ϕin the I=2 representation.
Write down the most general quartic potential V( ϕ) and study the possible symmetry breaking pattern.
IV .6. Anderson-Higgs Mechanism | 269
IV .6.5 Complete the derivation of the Feynman rules for the theory in (3) and compute the amplitude for the
physical process χ+χ→B+B.
IV .6.6 Derive (14). [Hint: The procedure is exactly the same as that used to obtain (III.4.9).] Write L=1
2AμQμνAν
withQμν=(∂2+M2)gμν−[1−(1/ξ)]∂μ∂νor in momentum space Qμν=−(k2−M2)gμν+[1−
(1/ξ)]kμkν. The propagator is the inverse of Qμν.
IV .6.7 Work out the ( ...)in(13)and the Feynman rules for the various interaction vertices.
IV .6.8 Using the Feynman rules derived in exercise IV .6.7 calculate the amplitude for the physical process ϕ/prime
1+
ϕ/prime
1→A+Aand show that the dependence on ξcancels out. Compare with the result in exercise IV .6.5.
[Hint: There are two diagrams, one with Aexchange and the other with ϕ2exchange.]
IV .6.9 Consider the theory defined in (12) with μ=0. Using the result of exercise IV .3.5 show that
Veff(ϕ)=1
4λϕ4+1
64π2(10λ2+3e4)ϕ4/parenleftbigg
logϕ2
M2−25
6/parenrightbigg
+... (16)
where ϕ2=ϕ2
1+ϕ2
2. This potential has a minimum away from ϕ=0 and thus the gauge symmetry
is spontaneously broken by quantum fluctuations. In chapter IV .3 we did not have the e4term and
argued that the minimum we got there was not to be trusted. But here we can balance the λϕ4against
e4ϕ4log(ϕ2/M2)forλof the same order of magnitude as e4. The minimum can be trusted. Show that
the spectrum of this theory consists of a massive scalar boson and a massive vector boson, with
m2(scalar)
m2(vector)=3
2πe2
4π(17)
For help, see S. Coleman and E. Weinberg, Phys. Rev. D7: 1888, 1973.
IV.7 Chiral Anomaly
Classical versus quantum symmetry
I have emphasized the importance of asking good questions. Here is another good one: Is
a symmetry of classical physics necessarily a symmetry of quantum physics?
We have a symmetry of classical physics if a transformation ϕ→ϕ+δϕleaves the action
S(ϕ) invariant. We have a symmetry of quantum physics if the transformation leaves the
path integral/integraltext
DϕeiS(ϕ)invariant.
When our question is phrased in this path integral language, the answer seems obvious:
Not necessarily. Indeed, the measure Dϕ may or may not be invariant.
Yet historically, field theorists took as almost self-evident the notion that any symmetry
of classical physics is necessarily a symmetry of quantum physics, and indeed, almostall the symmetries they encountered in the early days of field theory had the property ofbeing symmetries of both classical and quantum physics. For instance, we certainly expectquantum mechanics to be rotational invariant. It would be very odd indeed if quantumfluctuations were to favor a particular direction.
You have to appreciate the frame of mind that field theorists operated in to understand
their shock when they discovered in the late 1960s that quantum fluctuations can indeedbreak classical symmetries. Indeed, they were so shocked as to give this phenomenon therather misleading name “anomaly,” as if it were some kind of sickness of field theory. Withthe benefits of hindsight, we now understand the anomaly as being no less conceptuallyinnocuous as the elementary fact that when we change integration variables in an integralwe better not forget the Jacobian.
With the passing of time, field theorists have developed many different ways of looking
at the all important subject of anomaly. They are all instructive and shed different lights
on how the anomaly comes about. For this introductory text I choose to show the existenceof anomaly by an explicit Feynman diagram calculation. The diagram method is certainlymore laborious and less slick than other methods, but the advantage is that you will see
IV .7. Chiral Anomaly | 271
a classical symmetry vanishing in front of your very eyes! No smooth formal argument
for us.
The lesser of two evils
Consider the theory of a single massless fermion L=¯ψiγμ∂μψ. You can hardly ask
for a simpler theory! Recall from chapter II.1 that Lis manifestly invariant under the
separate transformations ψ→eiθψandψ→eiθγ5ψ, corresponding to the conserved
vector current Jμ=¯ψγμψand the conserved axial current Jμ
5=¯ψγμγ5ψrespectively.
You should verify that ∂μJμ=0 and ∂μJμ
5=0 follow immediately from the classical
equation of motion iγμ∂μψ=0.
Let us now calculate the amplitude for a spacetime history in which a fermion-anti-
fermion pair is created at x1and another such pair is created at x2by the vector current,
with the fermion from one pair annihilating the antifermion from the other pair and theremaining fermion-antifermion pair being subsequently annihilated by the axial current.This is a long-winded way of describing the amplitude /angbracketleft0|TJ
λ
5(0)Jμ(x1)Jν(x2)|0/angbracketrightin
words, but I want to make sure that you know what I am talking about. Feynman tellsus that the Fourier transform of this amplitude is given by the two “triangle” diagrams infigure IV .7.1a and b.
/Delta1λμν(k1,k2)=(−1)i3/integraldisplayd4p
(2π)4
tr/parenleftbigg
γλγ5 1
/negationslashp− /negationslashqγν 1
/negationslashp− /negationslashk1γμ1
/negationslashp+γλγ5 1
/negationslashp− /negationslashqγμ 1
/negationslashp− /negationslashk2γν1
/negationslashp/parenrightbigg
(1)
withq=k1+k2. Note that the two terms are required by Bose statistics. The overall factor
of(−1)comes from the closed fermion loop.
Classically, we have two symmetries implying ∂μJμ=0 and ∂μJμ
5=0. In the quantum
theory, if ∂μJμ=0 continues to hold, then we should have k1μ/Delta1λμν=0 and k2ν/Delta1λμν=0,
and if ∂μJμ
5=0 continues to hold, then qλ/Delta1λμν=0. Now that we have things all set up,
we merely have to calculate /Delta1λμνto see if the two symmetries hold up under quantum
fluctuations. No big deal.
p p − q
p − k1γλγ5
γμγν
(a)p p − qγλγ5
(b)γμγν
p − k2
Figure IV .7.1
272 | IV . Symmetry and Symmetry Breaking
Before we blindly calculate, however, let us ask ourselves how sad we would be if either of
the two currents JμandJμ
5fails to be conserved. Well, we would be very upset if the vector
current is not conserved. The corresponding charge Q=/integraltext
d3xJ0counts the number of
fermions. We wouldn’t want our fermions to disappear into thin air or pop out of nowhere.Furthermore, it may please us to couple the photon to the fermion field ψ. In that case,
you would recall from chapter II.7 that we need ∂
μJμ=0 to prove gauge invariance and
hence show that the photon has only two degrees of polarization. More explicitly, imaginea photon line coming into the vertex labeled by μin figure IV .7.1a and b with propagator
(i/k
2
1)[ξ(k 1μk1ρ/k2
1)−gμρ]. The gauge dependent term ξ(k 1μk1ρ/k2
1)would not go away if
the vector current is not conserved, that is, if k1μ/Delta1λμνfails to vanish.
On the other hand, quite frankly, just between us friends, we won’t get too upset if
quantum fluctuation violates axial current conservation. Who cares if the axial chargeQ
5=/integraltext
d3xJ0
5is not constant in time?
Shifting integration variable
So, do k1μ/Delta1λμνandk2ν/Delta1λμνvanish? We will look over Professor Confusio’s shoulders as
he calculates k1μ/Delta1λμν. (We are now in the 1960s, long after the development of renor-
malization theory as described in chapter III.1 and Confusio has managed to get a tenuretrack assistant professorship.) He hits /Delta1
λμνas written in (1) with k1μand using what he
learned in chapter II.7 writes /negationslashk1in the first term as /negationslashp−(/negationslashp− /negationslashk1)and in the second term
as(/negationslashp− /negationslashk2)−(/negationslashp− /negationslashq), thus obtaining
k1μ/Delta1λμν(k1,k2)
=i/integraldisplayd4p
(2π)4tr(γλγ5 1
/negationslashp− /negationslashqγν 1
/negationslashp− /negationslashk1−γλγ5 1
/negationslashp− /negationslashk2γν1
/negationslashp) (2)
Just as in chapter II.7, Confusio recognizes that in the integrand the first term is just the
second term with the shift of the integration variable p→p−k1. The two terms cancel
and Professor Confusio publishes a paper saying k1μ/Delta1λμν=0, as we all expect.
Remember back in chapter II.7 I said we were going to worry later about whether it is
legitimate to shift integration variables. Now is the time to worry!
You could have asked your calculus teacher long ago when it is legitimate to shift
integration variables. When is/integraltext+∞
−∞dpf (p +a)equal to/integraltext+∞
−∞dpf (p) ? The difference
between these two integrals is
/integraldisplay+∞
−∞dp(ad
dpf( p)+...)=a(f(+∞)−f(−∞)) +...
Clearly, if f(+∞) andf(−∞) are two different constants, then it is not okay to shift. But
if the integral/integraltext+∞
−∞dpf (p) is convergent, or even logarithmically divergent, it is certainly
okay. It was okay in chapter II.7 but definitely not here in (2)!
IV .7. Chiral Anomaly | 273
As usual, we rotate the Feynman integrand to Euclidean space. Generalizing our obser-
vation above to d-dimensional Euclidean space, we have
/integraldisplay
dd
Ep[f( p+a)−f( p) ]=/integraldisplay
dd
Ep[aμ∂μf( p)+...]
which by Gauss’s theorem is given by a surface integral over an infinitely large sphere
enclosing all of Euclidean spacetime and hence equal to
lim
P→∞aμ/parenleftbiggPμ
P/parenrightbigg
f( P) Sd−1(P)
where Sd−1(P) is the area of a (d−1)-dimensional sphere (see appendix D) and where an
average over the surface of the sphere is understood. (Recall from our experience evaluatingFeynman diagrams that the average of P
μPν/P2is equal to1
4ημνby a symmetry argument,
with the normalization1
4fixed by contracting with ημν.)Rotating back, we have for a 4-
dimensional Minkowskian integral
/integraldisplay
d4p[f( p+a)−f( p) ]=lim
P→∞iaμ/parenleftbiggPμ
P/parenrightbigg
f( P) ( 2π2P3) (3)
Note the ifrom Wick rotating back.
Applying (3) with
f( p)=tr/parenleftbigg
γλγ5 1
/negationslashp− /negationslashk2γν1
/negationslashp/parenrightbigg
=tr[γ5(/negationslashp− /negationslashk2)γν/negationslashpγλ]
(p−k2)2p2=4iετνσλk2τpσ
(p−k2)2p2
we obtain
k1μ/Delta1λμν=i
(2π)4lim
P→∞i(−k 1)μPμ
P4iετνσλk2τPσ
P42π2P3=i
8π2ελντσk1τk2σ
Contrary to what Confusio said, k1μ/Delta1λμν/negationslash=0.
As I have already said, this would be a disaster. Fermion number is not conserved and
matter would be disintegrating all around us! What is the way out?
In fact, we are only marginally smarter than Professor Confusio. We did not notice that
the integral defining /Delta1λμνin (1) is linearly divergent and is thus not well defined.
Oops, even before we worry about calculating k1μ/Delta1λμνandk2ν/Delta1λμνwe better worry
about whether or not /Delta1λμνdepends on the physicist doing the calculation. In other
words, suppose another physicist chooses1to shift the integration variable pin the linearly
divergent integral in (1) by an arbitrary 4-vector aand define
/Delta1λμν(a,k1,k2)
=(−1)i3/integraldisplayd4p
(2π)4tr(γλγ5 1
/negationslashp+ /negationslasha− /negationslashqγν 1
/negationslashp+ /negationslasha− /negationslashk1γμ 1
/negationslashp+ /negationslasha)
+{μ,k1↔ν,k2} (4)
There can be as many results for the Feynman diagrams in figure IV .7.1a and b as there
are physicists! That would be the end of physics, or at least quantum field theory, for sure.
1This is the freedom of choice in labeling internal momenta mentioned in chapter I.7.
274 | IV . Symmetry and Symmetry Breaking
Well, whose result should we declare to be correct?
The only sensible answer is that we trust the person who chooses an asuch that
k1μ/Delta1λμν(a,k1,k2)andk2ν/Delta1λμν(a,k1,k2)vanish, so that the photon will have the right
number of degrees of freedom should we introduce a photon into the theory.
Let us compute /Delta1λμν(a,k1,k2)−/Delta1λμν(k1,k2)by applying (3) to f( p)=
tr(γλγ51
/negationslashp−/negationslashqγν 1
/negationslashp−/negationslashk1γμ1
/negationslashp). Noting that
f( P)=lim
P→∞tr(γλγ5/negationslashPγν/negationslashPγμ/negationslashP)
P6
=2Pμtr(γλγ5/negationslashPγν/negationslashP)−P2tr(γλγ5/negationslashPγνγμ)
P6=+4iP2Pσεσνμλ
P6
we see that
/Delta1λμν(a,k1,k2)−/Delta1λμν(k1,k2)=4i
8π2lim
P→∞aωPωPσ
P2εσνμλ+{μ,k1↔ν,k2}
=i
8π2εσνμλaσ+{μ,k1↔ν,k2} (5)
There are two independent momenta k1andk2in the problem, so we can take a=
α(k1+k2)+β(k1−k2). Plugging into (5), we obtain
/Delta1λμν(a,k1,k2)=/Delta1λμν(k1,k2)+iβ
4π2ελμνσ(k1−k2)σ (6)
Note that αdrops out.
As expected, /Delta1λμν(a,k1,k2)depends on β, and hence on a. Our unshakable desire to
have a conserved vector current, that is, k1μ/Delta1λμν(a,k1,k2)=0, now fixes the parameter β
upon recalling
k1μ/Delta1λμν(k1,k2)=i
8π2ελντσk1τk2σ
Hence, we must choose to deal with /Delta1λμν(a,k1,k2)withβ=−1
2.
One way of viewing all this is to say that the Feynman rules do not suffice in determining
/angbracketleft0|TJλ
5(0)Jμ(x1)Jν(x2)|0/angbracketright. They have to be supplemented by vector current conservation.
The amplitude /angbracketleft0|TJλ
5(0)Jμ(x1)Jν(x2)|0/angbracketrightis defined by /Delta1λμν(a,k1,k2)withβ=−1
2.
Quantum fluctuation violates axial current conservation
Now we come to the punchline of the story. We insisted that the vector current be conserved.
Is the axial current also conserved?
T o answer this question, we merely have to compute
qλ/Delta1λμν(a,k1,k2)=qλ/Delta1λμν(k1,k2)+i
4π2εμνλσk1λk2σ (7)
IV .7. Chiral Anomaly | 275
By now, you know how to do this:
qλ/Delta1λμν(k1,k2)=i/integraldisplayd4p
(2π)4tr/parenleftbigg
γ5 1
/negationslashp− /negationslashqγν 1
/negationslashp− /negationslashk1γμ
−γ5 1
/negationslashp− /negationslashk2γν1
/negationslashpγμ/parenrightbigg
+{μ,k1↔ν,k2}
=i
4π2εμνλσk1λk2σ (8)
Indeed, you recognize that the integration has already been done in (2). We finally obtain
qλ/Delta1λμν(a,k1,k2)=i
2π2εμνλσk1λk2σ (9)
The axial current is not conserved!
In summary, in the simple theory L=¯ψiγμ∂μψwhile the vector and axial currents are
both conserved classically, quantum fluctuation destroys axial current conservation. Thisphenomenon is known variously as the anomaly, the axial anomaly, or the chiral anomaly.
Consequences of the anomaly
As I said, the anomaly is an extraordinarily rich subject. I will content myself with a seriesof remarks, the details of which you should work out as exercises.
1. Suppose we gauge our simple theory L=¯ψiγμ(∂μ−ieAμ)ψand speak of Aμas the photon
field. Then in figure IV .7.1 we can think of two photon lines coming out of the vertices labeledμandν. Our central result (9) can then be written elegantly as two operator equations:
Classical physics: ∂
μJμ
5=0 (10)
Qu antum physics: ∂μJμ
5=e2
(4π)2εμνλσFμνFλσ (11)
The divergence of the axial current ∂μJμ
5is not zero, but is an operator capable of producing
two photons.
2. Applying the same type of argument as in chapter IV .2 we can calculate the rate of the
decay π0→γ+γ. Indeed, historically people used the erroneous result (10) to deduce that
this experimentally observed decay cannot occur! See exercise IV .7.2. The resolution of thisapparent paradox led to the correct result (11).
3. Writing the Lagrangian in terms of left and right handed fields ψ
RandψLand introducing
the left and right handed currents Jμ
R≡¯ψRγμψRandJμ
L≡¯ψLγμψL, we can repackage the
anomaly as
∂μJμ
R=1
2e2
(4π)2εμνλσFμνFλσ
and
∂μJμ
L=−1
2e2
(4π)2εμνλσFμνFλσ (12)
276 | IV . Symmetry and Symmetry Breaking
(Hence the name chiral!) We can think of left handed and right handed fermions running
around the loop in figure IV .7.1, contributing oppositely to the anomaly.
4. Consider the theory L=¯ψ(iγμ∂μ−m)ψ . Then invariance under the transformation ψ→
eiθγ5ψis spoiled by the mass term. Classically, ∂μJμ
5=2m¯ψiγ5ψ: The axial current is
explicitly not conserved. The anomaly now says that quantum fluctuation produces anadditional term. In the theory L=¯ψ[iγ
μ(∂μ−ieAμ)−m]ψ, we have
∂μJμ
5=2m¯ψiγ5ψ+e2
(4π)2εμνλσFμνFλσ (13)
5. Recall that in chapter III.7 we introduced Pauli-Villars regulators to calculate vacuum
polarization. We subtract from the integrand what the integrand would have been if theelectron mass were replaced by some regulator mass. The analog of electron mass in (1) isin fact 0 and so we subtract from the integrand what the integrand would have been if 0were replaced by a regulator mass M. In other words, we now define
/Delta1
λμν(k1,k2)=(−1)i3/integraldisplayd4p
(2π)4tr/parenleftbigg
γλγ5 1
/negationslashp− /negationslashqγν 1
/negationslashp− /negationslashk1γμ1
/negationslashp
−γλγ5 1
/negationslashp− /negationslashq−Mγν 1
/negationslashp− /negationslashk1−Mγμ 1
/negationslashp−M/parenrightbigg
+{μ,k1↔ν,k2}. (14)
Note that as p→∞ the integrand now vanishes faster than 1 /p3. This is in accordance
with the philosophy of regularization outlined in chapters III.1 and III.7: For p/lessmuchM, the
threshold of ignorance, the integrand is unchanged. But for p/greatermuchM, the integrand is cut
off. Now the integral in (14) is superficially logarithmically divergent and we can shift theintegration variable pat will.
So how does the chiral anomaly arise? By including the regulator mass Mwe have broken
axial current conservation explicitly. The anomaly is the statement that this breaking persistseven when we let Mtend to infinity. It is extremely instructive (see exercise IV .7.4) to work
this out.
6. Consider the nonabelian theory L=¯ψiγ
μ(∂μ−igAa
μTa)ψ. We merely have to include in
the Feynman amplitude a factor of Taat the vertex labeled by μand a factor of Tbat the vertex
labeled by ν. Everything goes through as before except that in summing over all the different
fermions that run around the loop we obtain a factor tr TaTb. Thus, we see instantly that
in a nonabelian gauge theory
∂μJμ
5=g2
(4π)2εμνλσtrFμνFλσ (15)
where Fμν=Fa
μνTais the matrix field strength defined in chapter IV .5. Nonabelian sym-
metry tells us something remarkable: The object εμνλσtrFμνFλσcontains not only a term
quadratic in A, but also terms cubic and quartic in A, and hence there is also a chiral
anomaly with three and four gauge bosons coming in, as indicated in figure IV .7.2a andb. Some people refer to the anomaly produced in figures IV .7.1 and IV .7.2 as the triangle,square, and pentagon anomaly. Historically, after the triangle anomaly was discovered, therewas a controversy as to whether the square and pentagon anomaly existed. The nonabelian
IV .7. Chiral Anomaly | 277
γλγ5
(a)γλγ5
(b)TcTa
Tb TcTbTaTd
Figure IV .7.2
w2
w1
Figure IV .7.3
symmetry argument given here makes things totally obvious, but at the time people calcu-
lated Feynman diagrams explicitly and, as we just saw, there are subtleties lying in wait forthe unwary.
7. We will see in chapter V .7 that the anomaly has deep connections to topology.8. We computed the chiral anomaly in the free theory L=¯ψ(iγ
μ∂μ−m)ψ . Suppose we
couple the fermion to a scalar field by adding fϕ¯ψψ or to the electromagnetic field for
that matter. Now we have to calculate higher order diagrams such as the three-loop diagramin figure IV .7.3. You would expect that the right-hand side of (9) would be multiplied by1+h(f ,e,...), where his some unknown function of all the couplings in the theory.
Surprise! Adler and Bardeen proved that h=0. This apparently miraculous fact, known as
the nonrenormalization of the anomaly, can be understood heuristically as follows. Beforewe integrate over the momenta of the scalar propagators in figure IV .7.3 (labeled by w
1
andw2)the Feynman integrand has seven fermion propagators and thus is more than
sufficiently convergent that we can shift integration variables with impunity. Thus, beforewe integrate over w
1andw2all the appropriate Ward identities are satisfied, for instance,
qλ/Delta1λμν
3 loops(k1,k2;w1,w2)=0. You can easily complete the proof. You will give a proof2based
on topology in exercise V .7.13.
2For a simple proof not involving topology, see J. Collins, Renormalization , p. 352.
278 | IV . Symmetry and Symmetry Breaking
(a)π°
γγ
(b)π°
γ γ
Figure IV .7.4
9. The preceding point was of great importance in the history of particle physics as it led directly
to the notion of color, as we will discuss in chapter VII.3. The nonrenormalization of theanomaly allowed the decay amplitude for π
0→γ+γto be calculated with confidence in the
late 1960s. In the quark model of the time, the amplitude is given by an infinite number ofFeynman diagrams, as indicated in figure IV .7.4 (with a quark running around the fermionloop), but the nonrenormalization of the anomaly tells us that only figure IV .7.4a contributes.In other words, the amplitude does not depend on the details of the strong interaction. Thatit came out a factor of 3 too small suggested that quarks come in 3 copies, as we will see inchapter VII.3.
10. It is natural to speculate as to whether quarks and leptons are composites of yet more
fundamental fermions known as preons. The nonrenormalization of the chiral anomalyprovides a powerful tool for this sort of theoretical speculation. No matter how complicatedthe relevant interactions might be, as long as they are described by field theory as we knowit, the anomaly at the preon level must be the same as the anomaly at the quark-lepton level.This so-called anomaly matching condition
3severely constrains the possible preon theories.
11. Historically, field theorists were deeply suspicious of the path integral, preferring the
canonical approach. When the chiral anomaly was discovered, some people even argued thatthe existence of the anomaly proved that the path integral was wrong. Look, these peoplesaid, the path integral
/integraldisplay
D¯ψDψ e
i/integraltext
d4x¯ψiγμ(∂μ−iAμ)ψ(16)
is too stupid to tell us that it is not invariant under the chiral transformation ψ→eiθγ5ψ.
Fujikawa resolved the controversy by showing that the path integral did know about theanomaly: Under the chiral transformation the measure D¯ψDψ changes by a Jacobian.
Recall that this was how I motivated this chapter: The action may be invariant but not thepath integral.
3G. ’t Hooft, in: G. ’t Hooft et al., eds., Recent Developments in Gauge Theories ; A. Zee, Phys. Lett. 95B:290, 1980.
IV .7. Chiral Anomaly | 279
Exercises
IV .7.1 Derive (11) from (9). The momentum factors k1λandk2σin (9) become the two derivatives in FμνFλσin
(11).
IV .7.2 Following the reasoning in chapter IV .2 and using the erroneous (10) show that the decay amplitude for
the decay π0→γ+γwould vanish in the ideal world in which the π0is massless. Since the π0does
decay and since our world is close to the ideal world, this provided the first indication historically that(10) cannot possibly be valid.
IV .7.3 Repeat all the calculations in the text for the theory L=¯ψ(iγ
μ∂μ−m)ψ .
IV .7.4 T ake the Pauli-Villars regulated /Delta1λμν(k1,k2)and contract it with qλ. The analog of the trick in chapter II.7
is to write /negationslashqγ5in the second term as [2 M+(/negationslashp−M)−(/negationslashp− /negationslashq+M)]γ5. Now you can freely shift
integration variables. Show that
qλ/Delta1λμν(k1,k2)=− 2M/Delta1μν(k1,k2) (17)
where
/Delta1μν(k1,k2)≡(−1)i3/integraldisplayd4p
(2π)4
tr/parenleftbigg
γ5 1
/negationslashp− /negationslashq−Mγν 1
/negationslashp− /negationslashk1−Mγμ 1
/negationslashp−M/parenrightbigg
+{μ,k1↔ν,k2}
Evaluate /Delta1μνand show that /Delta1μνgoes as 1 /M in the limit M→∞ and so the right hand side of (17)
goes to a finite limit. The anomaly is what the regulator leaves behind as it disappears from the lowenergy spectrum: It is like the smile of the Cheshire cat. [We can actually argue that /Delta1
μνgoes as 1 /M
without doing a detailed calculation. By Lorentz invariance and because of the presence of γ5,/Delta1μν
must be proportional to εμνλρk1λk2ρ, but by dimensional analysis, /Delta1μνmust be some constant times
εμνλρk1λk2ρ/M . You might ask why we can’t use something like 1 /(k2
1)1
2instead of 1 /M to make the
dimension come out right. The answer is that from your experience in evaluating Feynman diagrams in
(3+1)-dimensional spacetime you can never get a factor like 1 /(k2
1)1
2.]
IV .7.5 There are literally Nways of deriving the anomaly. Here is another. Evaluate
/Delta1λμν(k1,k2)=(−1)i3/integraldisplayd4p
(2π)4
tr/parenleftbigg
γλγ5 1
/negationslashp− /negationslashq−mγν 1
/negationslashp− /negationslashk1−mγμ 1
/negationslashp−m/parenrightbigg
+{μ,k1↔ν,k2}
in the massive fermion case not by brute force but by first using Lorentz invariance to write
/Delta1λμν(k1,k2)=ελμνσk1σA1+...+εμνστk1σk2τkλ
2A8
where Ai≡Ai(k2
1,k2
2,q2)are eight functions of the three Lorentz scalars in the problem. You are
supposed to fill in the dots. By counting powers as in chapters III.3 and III.7 show that two of thesefunctions are given by superficially logarithmically divergent integrals while the other six are given byperfectly convergent integrals. Next, impose Bose statistics and vector current conservation k
1μ/Delta1λμν=
0=k2ν/Delta1λμνto show that we can avoid calculating the superficially logarithmically divergent integrals.
Compute the convergent integrals and then evaluate qλ/Delta1λμν(k1,k2).
IV .7.6 Discuss the anomaly by studying the amplitude
/angbracketleft0|TJλ
5(0)Jμ
5(x1)Jν
5(x2)|0/angbracketright
280 | IV . Symmetry and Symmetry Breaking
given in lowest orders by triangle diagrams with axial currents at each vertex. [Hint: Call the momentum
space amplitude /Delta1λμν
5(k1,k2).] Show by using (γ5)2=1 and Bose symmetry that
/Delta1λμν
5(k1,k2)=1
3[/Delta1λμν(a,k1,k2)+/Delta1μνλ(a,k2,−q)+/Delta1νλμ(a,−q,k1)]
Now use (9) to evaluate qλ/Delta1λμν
5(k1,k2).
IV .7.7 Define the fermionic measure Dψ in (16) carefully by going to Euclidean space. Calculate the Jacobian
upon a chiral transformation and derive the anomaly. [Hint: For help, see K. Fujikawa, Phys. Rev. Lett.
42: 1195, 1979.]
IV .7.8 Compute the pentagon anomaly by Feynman diagrams in order to check remark 6 in the text. In other
words, determine the coefficient cin∂μJμ
5=...+cεμνλσtrAμAνAλAσ.
Part V Field Theory and Collective Phenomena
I mentioned in the introduction that one of the more intellectually satisfying developments
in the last two or three decades has been the increasingly important role played by fieldtheoretic methods in condensed matter physics. This is a rich and diverse subject; in thisand subsequent chapters I can barely describe the tip of the iceberg and will have to contentmyself with a few selected topics.
Historically, field theory was introduced into condensed matter physics in a rather direct
and straightforward fashion. The nonrelativistic electrons in a condensed matter systemcan be described by a field ψ, along the lines discussed in chapter III.5. Field theoretic
Lagrangians may then be written down, Feynman diagrams and rules developed, and soon and so forth. This is done in a number of specialized texts. What we present here is to alarge extent the more modern view of an effective field theoretic description of a condensedmatter system, valid at low energy and momentum. One of the fascinations of condensedmatter physics is that due to highly nontrivial many body effects the low energy degreesof freedom might be totally different from the electrons we started out with. A particularlystriking example (to be discussed in chapter VI.2) is the quantum Hall system, in whichthe low energy effective degree of freedom carries fractional charge and statistics.
Another advantage of devoting a considerable portion of a field theory book to condensed
matter physics is that historically and pedagogically it is much easier to understand therenormalization group in condensed matter physics than in particle physics.
I will defiantly not stick to a legalistic separation between condensed matter and particle
physics. Some of the topics treated in Parts V and VI actually belong to particle physics.And of course I cannot be responsible for explaining condensed matter physics, any morethan I could be responsible for explaining particle physics in chapter IV .2.
This page intentionally left blank
V.1 Superfluids
Repulsive bosons
Consider a finite density ¯ρof nonrelativistic bosons interacting with a short ranged repul-
sion. Return to (III.5.11):
L=iϕ†∂0ϕ−1
2m∂iϕ†∂iϕ−g2(ϕ†ϕ−¯ρ)2(1)
The last term is exactly the Mexican well potential of chapter IV .1, forcing the magnitude
ofϕto be close to√¯ρ, thus suggesting that we use polar variables ϕ≡√ρeiθas we did
in (III.5.7). Plugging in and dropping the total derivative (i/2)∂0ρ, we obtain
L=−ρ∂0θ−1
2m/bracketleftbigg1
4ρ(∂iρ)2+ρ(∂iθ)2/bracketrightbigg
−g2(ρ−¯ρ)2(2)
Spontaneous symmetry breaking
As in chapter IV .1 write√ρ=√¯ρ+h(the vacuum expectation value of ϕis√¯ρ), assume
h/lessmuch√¯ρ, and expand1:
L=− 2/radicalbig
¯ρh∂ 0θ−¯ρ
2m(∂iθ)2−1
2m(∂ih)2−4g2¯ρh2+... (3)
Picking out the terms up to quadratic in hin (3) we use the “central identity of quantum
field theory” (see appendix A) to integrate out h, obtaining
L=¯ρ∂0θ1
4g2¯ρ−(1/2m)∂2
i∂0θ−¯ρ
2m(∂iθ)2+...
=1
4g2(∂0θ)2−¯ρ
2m(∂iθ)2+... (4)
1Note that we have dropped the (potentially interesting) term −¯ρ∂0θbecause it is a total divergence.
284 | V . Field Theory and Collective Phenomena
In the second equality we assumed that we are looking at processes with wave number k
small compared to/radicalbig
8g2¯ρm so that (1/2m)∂2
iis negligible compared to 4 g2¯ρ. Thus, we see
that there exists in this fluid of bosons a gapless mode (often referred to as the phonon)with the dispersion
ω2=2g2¯ρ
m/vectork2(5)
The learned among you will have realized that we have obtained Bogoliubov’s classic result
without ever doing a Bogoliubov rotation.2
Let me briefly remind you of Landau’s idealized argument3that a linearly dispersing
mode (that is, ωis linear in k)implies superfluidity. Consider a mass Mof fluid flowing
down a tube with velocity v. It could lose momentum and slow down to velocity v/prime
by creating an excitation of momentum k:Mv=Mv/prime+/planckover2pik. This is only possible with
sufficient energy to spare if1
2Mv2≥1
2Mv/prime2+/planckover2piω(k) . Eliminating v/primewe obtain for M
macroscopic v≥ω/k . For a linearly dispersing mode this gives a critical velocity vc≡ω/k
below which the fluid cannot lose momentum and is hence super. [Thus, from (5) theidealized v
c=g√2¯ρ/m .]
Suitably scaling the distance variable, we can summarize the low energy physics of
superfluidity in the compact Lagrangian
L=1
4g2(∂μθ)2(6)
which we recognize as the massless version of the scalar field theory we studied in part I,
but with the important proviso that θis a phase angle field, that is, θ(x) andθ(x)+2πare
really the same. This gapless mode is evidently the Nambu-Goldstone boson associatedwith the spontaneous breaking of the global U(1)symmetry ϕ→e
iαϕ.
Linearly dispersing gapless mode
The physics here becomes particularly clear if we think about a gas of free bosons. We can
give a momentum /planckover2pi/vectorkto any given boson at the cost of only (/planckover2pi/vectork)2/2min energy. There
exist many low energy excitations in a free boson system. But as soon as a short rangedrepulsion is turned on between the bosons, a boson moving with momentum /vectorkwould
affect all the other bosons. A density wave is set up as a result, with energy proportionaltokas we have shown in (5). The gapless mode has gone from quadratically dispersing to
linearly dispersing. There are far fewer low energy excitations. Specifically, recall that thedensity of states is given by N(E) ∝k
D−1(dk/dE) . For example, for D=2 the density of
states goes from N(E) ∝constant (in the presence of quadratically dispersing modes) to
N(E) ∝E(in the presence of linearly dispersing modes) at low energies.
2L. D. Landau and E. M. Lifshitz, Statistical Physics , p. 238.
3Ibid., p. 192.
V .1. Superfluids | 285
As was emphasized by Feynman4among others, the physics of superfluidity lies not
in the presence of gapless excitations, but in the paucity of gapless excitations. (After all,the Fermi liquid has a continuum of gapless modes.) There are too few modes that thesuperfluid can lose energy and momentum to.
Relativistic versus nonrelativistic
This is a good place to discuss one subtle difference between spontaneous symmetrybreaking in relativistic and nonrelativistic theories. Consider the relativistic theory studiedin chapter IV .1: L=(∂/Phi1
†)(∂/Phi1) −λ(/Phi1†/Phi1−v2)2. It is often convenient to take the λ→∞
limit holding vfixed. In the language used in chapter IV .1 “climbing the wall” costs
infinitely more energy than “rolling along the gutter.” The resulting theory is defined by
L=(∂/Phi1†)(∂/Phi1) (7)
with the constraint /Phi1†/Phi1=v2. This is known as a nonlinear σmodel, about which much
more in chapter VI.4.
The existence of a Nambu-Goldstone boson is particularly easy to see in the nonlinear σ
model. The constraint is solved by /Phi1=veiθ, which when plugged into Lgives L=v2(∂θ)2.
There it is: the Nambu-Goldstone boson θ.
Let’s repeat this in the nonrelativistic domain. T ake the limit g2→∞ with¯ρheld fixed
so that (1) becomes
L=iϕ†∂0ϕ−1
2m∂iϕ†∂iϕ (8)
with the constraint ϕ†ϕ=¯ρ. But now if we plug the solution of the constraint ϕ=√¯ρeiθ
into L(and drop the total derivative −¯ρ∂0θ), we get L=−(¯ρ/2m)(∂iθ)2with the equation
of motion ∂2
iθ=0. Oops, what is this? It’s not even a propagating degree of freedom?
Where is the Nambu-Goldstone boson?
Knowing what I already told you, you are not going to be puzzled by this apparent
paradox5for long, but believe me, I have stumped quite a few excellent relativistic minds
with this one. The Nambu-Goldstone boson is still there, but as we can see from (5) itspropagation velocity ω/k scales to infinity as gand thus it disappears from the spectrum
for any nonzero /vectork.
Why is it that we are allowed to go to this “nonlinear” limit in the relativistic case?
Because we have Lorentz invariance! The velocity of a linearly dispersing mode, if such amode exists, is guaranteed to be equal to 1.
4R. P. Feynman, Statistical Mechanics .
5This apparent paradox was discussed by A. Zee, “From Semionics to T opological Fluids” in O. J. P. ´Ebolic et
al., eds., Particle Physics , p. 415.
286 | V . Field Theory and Collective Phenomena
Exercises
V .1.1 Verify that the approximation used to reach (3) is consistent.
V .1.2 T o confine the superfluid in an external potential W(/vectorx)we would add the term −W(/vectorx)ϕ†(/vectorx,t)ϕ(/vectorx,t)
to (1). Derive the corresponding equation of motion for ϕ. The equation, known as the Gross-Pitaevski
equation, has been much studied in recent years in connection with the Bose-Einstein condensate.
V.2Euclid, Boltzmann, Hawking, and
Field Theory at Finite T emperature
Statistical mechanics and Euclidean field theory
I mentioned in chapter I.2 that to define the path integral more rigorously we should
perform a Wick rotation t=−itE. The scalar field theory, instead of being defined by the
Minkowskian path integral
Z=/integraldisplay
Dϕe(i//planckover2pi)/integraltext
ddx[1
2(∂ϕ)2−V( ϕ) ](1)
is then defined by the Euclidean functional integral
Z=/integraldisplay
Dϕe−(1//planckover2pi)/integraltext
dd
Ex[1
2(∂ϕ)2+V( ϕ) ]=/integraldisplay
Dϕe−(1//planckover2pi)E(ϕ)(2)
where ddx=−idd
Ex, with dd
Ex≡dtEd(d−1)x.I n( 1 )(∂ϕ)2=(∂ϕ/∂t)2−(/vector∇ϕ)2, while in (2)
(∂ϕ)2=(∂ϕ/∂t E)2+(/vector∇ϕ)2: The notation is a tad confusing but I am trying not to introduce
too many extraneous symbols. You may or may not find it helpful to think of (/vector∇ϕ)2+V( ϕ)
as one unit, untouched by Wick rotation. I have introduced E(ϕ)≡/integraltext
dd
Ex[1
2(∂ϕ)2+V( ϕ) ],
which may naturally be regarded as a static energy functional of the field ϕ(x) . Thus,
given a configuration ϕ(x) ind-dimensional space, the more it varies, the less likely it is
to contribute to the Euclidean functional integral Z.
The Euclidean functional integral (2) may remind you of statistical mechanics. Indeed,
Herr Boltzmann taught us that in thermal equilibrium at temperature T=1/β, the
probability for a configuration to occur in a classical system or the probability for a state tooccur in a quantum system is just the Boltzmann factor e
−βEsuitably normalized, where E
is to be interpreted as the energy of the configuration in a classical system or as the energyeigenvalue of the state in a quantum system. In particular, recall the classical statisticalmechanics of an N-particle system for which
E(p ,q)=/summationdisplay
i1
2mp2
i+V( q 1,q2,... ,qN)
288 | V . Field Theory and Collective Phenomena
The partition function is given (up to some overall constant) by
Z=/productdisplay
i/integraldisplay
dpidqie−βE(p ,q)
After doing the integrals over pwe are left with the (reduced) partition function
Z=/productdisplay
i/integraldisplay
dqie−βV(q 1,q2,...,qN)
Promoting this to a field theory as in chapter I.3, letting i→xandqi→ϕ(x) as before, we
see that the partition function of a classical field theory with the static energy functionalE(ϕ) has precisely the form in (2), upon identifying the symbol /planckover2pias the temperature
T=1/β. Thus,
Euclidean quantum field theory in d-dimensional spacetime
∼Classical statistical mechanics in d-dimensional space(3)
Functional integral representation of the quantum partition function
More interestingly, we move on to quantum statistical mechanics. The integration over
phase space {p,q}is replaced by a trace, that is, a sum over states: Thus the partition
function of a quantum mechanical system (say of a single particle to be definite) with theHamiltonian His given by
Z=tre−βH=/summationdisplay
n/angbracketleftn|e−βH|n/angbracketright
In chapter I.2 we worked out the integral representation of /angbracketleftF|e−iHT|I/angbracketright. (You should
not confuse the time Twith the temperature Tof course.) Suppose we want an integral
representation of the partition function. No need to do any more work! We simply replacethe time Tby−iβ , set|I/angbracketright=|F/angbracketright=|n/angbracketrightand sum over |n/angbracketrightto obtain
Z=tre−βH=/integraldisplay
PBCDqe−/integraltextβ
0dτL(q)(4)
T racing the steps from (I.2.3) to (I.2.5) you can verify that here L(q)=1
2(dq/dτ)2+V( q)
is precisely the Lagrangian corresponding to Hin the Euclidean time τ. The integral over
τruns from 0 to β. The trace operation sets the initial and final states equal and so the
functional integral should be done over all paths q(τ) with the boundary condition q(0)=
q(β) . The subscript PBC reminds us of this all important periodic boundary condition.
The extension to field theory is immediate. If His the Hamiltonian of a quantum field
theory in D-dimensional space [and hence d=(D+1)-dimensional spacetime], then the
partition function (4) is
Z=tre−βH=/integraldisplay
PBCDϕe−/integraltextβ
0dτ/integraltext
dDxL(ϕ)(5)
V .2. Finite T emperatures | 289
with the integral evaluated over all paths ϕ(/vectorx,τ)such that
ϕ(/vectorx,0)=ϕ(/vectorx,β) (6)
(Here ϕrepresents all the Bose fields in the theory.)
A remarkable result indeed! T o study a field theory at finite temperature all we have to
do is rotate it to Euclidean space and impose the boundary condition (6). Thus,
Euclidean quantum field theory in (D+1)-dimensional
spacetime, 0 ≤τ<β
∼Quantum statistical mechanics in D-dimensional space(7)
In the zero temperature limit β→∞ we recover from (5) the standard Wick-rotated
quantum field theory over an infinite spacetime, as we should.
Surely you would hit it big with mystical types if you were to tell them that temperature
is equivalent to cyclic imaginary time. At the arithmetic level this connection comes merelyfrom the fact that the central objects in quantum physics e
−iHTand in thermal physics
e−βHare formally related by analytic continuation. Some physicists, myself included, feel
that there may be something profound here that we have not quite understood.
Finite temperature Feynman diagrams
If we so desire, we can develop the finite temperature perturbation theory of (5), workingout the Feynman rules and so forth. Everything goes through as before with one majordifference stemming from the condition (6) ϕ(/vectorx,τ=0)=ϕ(/vectorx,τ=β). Clearly, when
we Fourier transform with the factor e
iωτ, the Euclidean frequency ωcan take on only
discrete values ωn≡(2π/β)n , withnan integer. The propagator of the scalar field becomes
1/(k2
4+/vectork2)→1/(ω2
n+/vectork2). Thus, to evaluate the partition function, we simply write the
relevant Euclidean Feynman diagrams and instead of integrating over frequency we sumover a discrete set of frequencies ω
n=(2πT)n ,n=− ∞ ,... ,+∞ . In other words, after
you beat a Feynman integral down to the form/integraltext
dd
EkF(k2
E), all you have to do is replace it
by 2πT/summationtext
n/integraltext
dDkF[(2πT)2n2+/vectork2].
It is instructive to see what happens in the high-temperature T→∞ limit. In summing
overωn, then=0 term dominates since the combination (2πT)2n2+/vectork2occurs in the
denominator. Hence, the diagrams are evaluated effectively in D-dimensional space. We
lose a dimension! Thus,
Euclidean quantum field theory in D-dimensional spacetime
∼High-temperature quantum statistical mechanics in
D-dimensional space(8)
This is just the statement that at high-temperature quantum statistical mechanics goes
classical [compare (3)].
290 | V . Field Theory and Collective Phenomena
An important application of quantum field theory at finite temperature is to cosmology:
The early universe may be described as a soup of elementary particles at some hightemperature.
Hawking radiation
Hawking radiation from black holes is surely the most striking prediction of gravitationalphysics of the last few decades. The notion of black holes goes all the way back to Michelland Laplace, who noted that the escape velocity from a sufficiently massive object mayexceed the speed of light. Classically, things fall into black holes and that’s that. But withquantum physics a black hole can in fact radiate like a black body at a characteristictemperature T.
Remarkably, with what little we learned in chapter I.11 and here, we can actually
determine the black hole temperature. I hasten to add that a systematic development wouldbe quite involved and fraught with subtleties; indeed, entire books are devoted to thissubject. However, what we need to do is more or less clear. Starting with chapter I.11, wewould have to develop quantum field theory (for instance, that of a scalar field ϕ)in curved
spacetime, in particular in the presence of a black hole, and ask what a vacuum state (i.e.,a state devoid of ϕquanta) in the far past evolves into in the far future. We would find a
state filled with a thermal distribution of ϕquanta. We will not do this here.
In hindsight, people have given numerous heuristic arguments for Hawking radiation.
Here is one. Let us look at the Schwarzschild solution (see chapter I.11)
ds2=/parenleftbigg
1−2GM
r/parenrightbigg
dt2−/parenleftbigg
1−2GM
r/parenrightbigg−1
dr2−r2dθ2−r2sin2θd φ2(9)
At the horizon r=2GM , the coefficients of dt2anddr2change sign, indicating that
time and space, and hence energy and momentum, are interchanged. Clearly, somethingstrange must occur. With quantum fluctuations, particle and antiparticle pairs are alwayspopping in and out of the vacuum, but normally, as we had discussed earlier, the uncer-tainty principle limits the amount of time /Delta1tthe pairs can exist to ∼1//Delta1E . Near the black
hole horizon, the situation is different. A pair can fluctuate out of the vacuum right at thehorizon, with the particle just outside the horizon and the antiparticle just inside; heuris-tically the Heisenberg restriction on /Delta1t may be evaded since what is meant by energy
changes as we cross the horizon. The antiparticle falls in while the particle escapes to spa-tial infinity. Of course, a hand-waving argument like this has to be backed up by detailedcalculations.
If black holes do indeed radiate at a definite temperature T, and that is far from obvious
a priori, we can estimate Teasily by dimensional analysis. From (9) we see that only the
combination GM , which evidently has the dimension of a length, can come in. Since T
has the dimension of mass, that is, length inverse, we can only have T∝1/GM .
T o determine Tprecisely, we resort to a rather slick argument. I warn you from the
outset that the argument will be slick and should be taken with a grain of salt. It is onlymeant to whet your appetite for a more correct treatment.
V .2. Finite T emperatures | 291
Imagine quantizing a scalar field theory in the Schwarzschild metric, along the line
described in chapter I.11. If upon Wick rotation the field “feels” that time is periodic withperiod β, then according to what we have learned in this chapter the quanta of the scalar
field would think that they are living in a heat bath with temperature T=1/β.
Setting t→−iτ, we rotate the metric to
ds2=−/bracketleftBigg/parenleftbigg
1−2GM
r/parenrightbigg
dτ2+/parenleftbigg
1−2GM
r/parenrightbigg−1
dr2+r2dθ2+r2sin2θdφ2/bracketrightBigg
(10)
In the region just outside the horizon r>∼2GM , we perform the general coordinate trans-
formation (τ,r)→(α,R)so that the first two terms in ds2become R2dα2+dR2, namely
the length element squared of flat 2-dimensional Euclidean space in polar coordinates.
T o leading order, we can write the Schwarzchild factor (1−2GM/r) as
(r−2GM)/(2 GM)≡γ2R2with the constant γto be determined. Then the second term
becomes dr2/(γ2R2)=(4GM)2γ2dR2, and thus we set γ=1/(4GM) to get the desired
dR2. The first two terms in −ds2are then given by R2(dτ/(4 GM))2+dR2. Thus the Eu-
clidean time is related to the polar angle by τ=4GMα and so has a period of 8πGM =β.
We obtain thus the Hawking temperature
T=1
8πGM=/planckover2pic3
8πGM(11)
Restoring /planckover2piby dimensional analysis, we see that Hawking radiation is indeed a quantum
effect.
It is interesting to note that the Wick rotated geometry just outside the horizon is given
by the direct product of a plane with a 2-sphere of radius 2 GM , although, this observation
is not needed for the calculation we just did.
Exercises
V .2.1 Study the free field theory L=1
2(∂ϕ)2−1
2m2ϕ2at finite temperature and derive the Bose-Einstein
distribution.
V .2.2 It probably does not surprise you that for fermionic fields the periodic boundary condition (6) is replaced
by an antiperiodic boundary condition ψ(/vectorx,0)=−ψ(/vectorx,β)in order to reproduce the results of chap-
ter II.5. Prove this by looking at the simplest fermionic functional integral. [Hint: The clearest expositionof this satisfying fact may be found in appendix A of R. Dashen, B. Hasslacher, and A. Neveu, Phys. Rev.
D12: 2443, 1975.]
V .2.3 It is interesting to consider quantum field theory at finite density, as may occur in dense astrophysical
objects or in heavy ion collisions. (In the previous chapter we studied a system of bosons at finite densityand zero temperature.) In statistical mechanics we learned to go from the partition function to the grand
partition function Z=tre
−β(H −μN), where a chemical potential μis introduced for every conserved
particle number N. For example, for noninteracting relativistic fermions, the Lagrangian is modified
toL=¯ψ(i/negationslash∂−m)ψ+μ¯ψγ0ψ. Note that finite density, as well as finite temperature, breaks Lorentz
invariance. Develop the subject of quantum field theory at finite density as far as you can.
V.3 Landau-Ginzburg Theory of Critical Phenomena
The emergence of nonanalyticity
Historically, the notion of spontaneous symmetry breaking, originating in the work of
Landau and Ginzburg on second-order phase transitions, came into particle physics fromcondensed matter physics.
Consider a ferromagnetic material in thermal equilibrium at temperature T. The mag-
netization /vectorM(x) is defined as the average of the atomic magnetic moments taken over
a region of a size much larger than the length scale characteristic of the relevant micro-scopic physics. (In this chapter, we are discussing a nonrelativistic theory and xdenotes the
spatial coordinates only.) We know that at low temperatures, rotational invariance is spon-taneously broken and that the material exhibits a bulk magnetization pointing in somedirection. As the temperature is raised past some critical temperature T
cthe bulk mag-
netization suddenly disappears. We understand that with increased thermal agitation theatomic magnetic moments point in increasingly random directions, canceling each otherout. More precisely, it was found experimentally that just below T
cthe magnetization |/vectorM|
vanishes as ∼(Tc−T)β, where the so-called critical exponent β/similarequal0.37.
This sudden change is known as a second order phase transition, an example of a
critical phenomenon. Historically, critical phenomena presented a challenge to theoreticalphysicists. In principle, we are to compute the partition function Z=tre
−H/Twith
the microscopic Hamiltonian H, but Zis apparently smooth in Texcept possibly at
T=0. Some physicists went as far as saying that nonanalytic behavior such as (Tc−
T)βis impossible and that within experimental error |/vectorM|actually vanishes as a smooth
function of T. Part of the importance of Onsager’s famous exact solution in 1944 of the 2-
dimensional Ising model is that it settled this question definitively. The secret is that aninfinite sum of terms each of which may be analytic in some variable need not be analyticin that variable. The trace in tr e
−H/Tsums over an infinite number of terms.
V .3. Theory of Critical Phenomena | 293
Arguing from symmetry
In most situations, it is essentially impossible to calculate Zstarting with the microscopic
Hamiltonian. Landau and Ginzburg had the brilliant insight that the form of the freeenergy Gas a function of /vectorMfor a system with volume Vcould be argued from general
principles. First, for /vectorMconstant in x, we have by rotational invariance
G=V[a/vectorM2+b(/vectorM2)2+...] (1)
where a,b,... are unknown (but expected to be smooth) functions of T. Landau and
Ginzburg supposed that avanishes at some temperature Tc. Unless there is some special
reason, we would expect that for TnearTcwe have a=a1(T−Tc)+...[rather than, say,
a=a2(T−Tc)2+...]. But you already learned in chapter IV .1 what would happen. For
T> T c,Gis minimized at /vectorM=0, but as Tdrops below Tc, new minima suddenly develop
at|/vectorM|=/radicalbig
(−a/ 2b)∼(Tc−T)1
2. Rotational symmetry is spontaneously broken, and the
mysterious nonanalytic behavior pops out easily.
T o include the possibility of /vectorMvarying in space, Landau and Ginzburg argued that G
must have the form
G=/integraldisplay
d3x{∂i/vectorM∂i/vectorM+a/vectorM2+b(/vectorM2)2+...} (2)
where the coefficient of the ( ∂i/vectorM)2term has been set to 1 by rescaling /vectorM. You would
recognize (2) as the Euclidean version of the scalar field theory we have been studying. Bydimensional analysis we see that 1 /√
asets the length scale. More precisely, for T> Tc,
let us turn on a perturbing external magnetic field /vectorH(x) by adding the term −/vectorH./vectorM.
Assuming /vectorMsmall and minimizing Gwe obtain (−∂2+a)/vectorM/similarequal/vectorH, with the solution
/vectorM(x) =/integraldisplay
d3y/integraldisplayd3k
(2π)3ei/vectork.(/vectorx−/vectory)
/vectork2+a/vectorH(y)
=/integraldisplay
d3y1
4π|/vectorx−/vectory|e−√a|/vectorx−/vectory|/vectorH(y) (3)
[Recall that we did the integral in (I.4.7)—admire the unity of physics!]
It is standard to define a correlation function </vectorM(x) /vectorM(0)> by asking what the mag-
netization /vectorM(x) will be if we use a magnetic field sharply localized at the origin to create
a magnetization /vectorM(0)there. We expect the correlation function to die off as e−|/vectorx|/ξover
some correlation length ξthat goes to infinity as Tapproaches Tcfrom above. The critical
exponent νis traditionally defined by ξ∼1/(T−Tc)ν.
In Landau-Ginzburg theory, also known as mean field theory, we obtain ξ=1/√aand
hence ν=1
2.
The important point is not how well the predicted critical exponents such as βandν
agree with experiment but how easily they emerge from Landau-Ginzburg theory. The
theory provides a starting point for a complete theory of critical phenomena, which waseventually developed by Kadanoff, Fisher, Wilson, and others using the renormalizationgroup (to be discussed in chapter VI.8).
294 | V . Field Theory and Collective Phenomena
The story goes that Landau had a logarithmic scale with which he ranked theoretical
physicists, with Einstein on top, and that after working out Landau-Ginzburg theory hemoved himself up by half a notch.
Exercise
V .3.1 Another important critical exponent γis defined by saying that the susceptibility χ≡(∂M/∂H) |H=0
diverges ∼1/|T−Tc|γasTapproaches Tc. Determine γin Landau-Ginzburg theory. [Hint: Instructively,
there are two ways of doing it: (a) Add −/vectorH./vectorMto (1) for /vectorMand/vectorHconstant in space and solve for /vectorM(/vectorH).
(b) Calculate the susceptibility function χij(x−y)≡[∂Mi(x)/∂H j(y)]|H=0and integrate over space.]
V.4 Superconductivity
Pairing and condensation
When certain materials are cooled below a certain critical temperature Tc, they suddenly be-
come superconducting. Historically, physicists had long suspected that the superconduct-ing transition, just like the superfluid transition, has something to do with Bose-Einsteincondensation. But electrons are fermions, not bosons, and thus they first have to pairinto bosons, which then condense. We now know that this general picture is substantiallycorrect: Electrons form Cooper pairs, whose condensation is responsible for superconduc-tivity.
With brilliant insight, Landau and Ginzburg realized that without having to know the
detailed mechanism driving the pairing of electrons into bosons, they could understanda great deal about superconductivity by studying the field ϕ(x) associated with these con-
densing bosons. In analogy with the ferromagnetic transition in which the magnetization
/vectorM(x) in a ferromagnet suddenly changes from zero to a nonzero value when the temper-
ature drops below some critical temperature, they proposed that ϕ(x) becomes nonzero
for temperatures below T
c. (In this chapter xdenotes spatial coordinates only.) In statisti-
cal physics, quantities such as /vectorM(x) andϕ(x) that change through a phase transition are
known as order parameters.
The field ϕ(x) carries two units of electric charge and is therefore complex. The dis-
cussion now unfolds much as in chapter V .3 except that ∂iϕshould be replaced by Diϕ≡
(∂i−i2eAi)ϕsinceϕis charged. Following Landau and Ginzburg and including the energy
of the external magnetic field, we write the free energy as
F=1
4F2
ij+|Diϕ|2+a|ϕ|2+b
2|ϕ|4+... (1)
which is clearly invariant under the U(1)gauge transformation ϕ→ei2e/Lambda1ϕandAi→Ai+
∂i/Lambda1. As before, setting the coefficient of |Diϕ|2equal to 1 just amounts to a normalization
choice for ϕ.
The similarity between (1) and (IV .6.1) should be evident.
296 | V . Field Theory and Collective Phenomena
Meissner effect
A hallmark of superconductivity is the Meissner effect, in which an external magnetic
field/vectorBpermeating the material is expelled from it as the temperature drops below Tc. This
indicates that a constant magnetic field inside the material is not favored energetically. Theeffective laws of electromagnetism in the material must somehow change at T
c. Normally,
a constant magnetic field would cost an energy of the order ∼/vectorB2V, where Vis the volume
of the material. Suppose that the energy density is changed from the standard /vectorB2to/vectorA2
(where as usual /vector∇×/vectorA=/vectorB). For a constant magnetic field /vectorB,/vectorAgrows as the distance and
hence the total energy would grow faster than V. After the material goes superconducting,
we have to pay an unacceptably large amount of extra energy to maintain the constantmagnetic field and so it is more favorable to expel the magnetic field.
Note that a term like /vectorA
2in the effective energy density preserves rotational and transla-
tional invariance but violates electromagnetic gauge invariance. But we already know howto break gauge invariance from chapter IV .6. Indeed, the U(1)gauge theory described there
and the theory of superconductivity described here are essentially the same, related by aWick rotation.
As in chapter V .3 we suppose that for temperature T/similarequalT
c,a/similarequala1(T−Tc)while b
remains positive. The free energy Fis minimized by ϕ=0 above Tc, and by |ϕ|=√−a/b ≡vbelow Tc. All this is old hat to you, who have learned that upon symmetry
breaking in a gauge theory the gauge field gains a mass. We simply read off from (1) that
F=1
4F2
ij+(2ev)2A2
i+... (2)
which is precisely what we need to explain the Meissner effect.
London penetration length and coherence length
Physically, the magnetic field does not drop precipitously from some nonzero value outside
the superconductor to zero inside, but drops over some characteristic length scale, calledthe London penetration length. The magnetic field leaks into the superconductor a bit overa length scale l, determined by the competition between the energy in the magnetic field
F
2
ij∼(∂A)2∼A2/l2and the Meissner term (2ev)2A2in (2). Thus, Landau and Ginzburg
obtained the London penetration length lL∼(1/ev)=(1/e)√b/−a.
Similarly, the characteristic length scale over which the order parameter ϕvaries is
known as the coherence length lϕ, which can be estimated by balancing the second and
third terms in (1), roughly (∂ϕ)2∼ϕ2/l2
ϕandaϕ2against each other, giving a coherence
length of order lϕ∼1/√−a.
Putting things together, we have
lL
lϕ∼√
b
e(3)
V .4. Superconductivity | 297
You might recognize from chapter IV .6 that this is just the ratio of the mass of the scalar
field to the mass of the vector field.
As I remarked earlier, the concept of spontaneous symmetry breaking went from con-
densed matter physics to particle physics. After hearing a talk at the University of Chicagoon the Bardeen-Cooper-Schrieffer theory of superconductivity by the young Schrieffer,Nambu played an influential role in bringing spontaneous symmetry breaking to the par-ticle physics community.
Exercises
V .4.1 Vary (1) to obtain the equation for Aand determine the London penetration length more carefully.
V .4.2 Determine the coherence length more carefully.
V.5 Peierls Instability
Noninteracting hopping electrons
The appearance of the Dirac equation and a relativistic field theory in a solid would be
surprising indeed, but yes, it is possible.
Consider the Hamiltonian
H=−t/summationdisplay
j(c†
j+1cj+c†
jcj+1) (1)
describing noninteracting electrons hopping on a 1-dimensional lattice (figure V .5.1). Here
cjannihilates an electron on site j. Thus, the first term describes an electron hopping from
sitejto site j+1 with amplitude t. We have suppressed the spin labels. This is just about
the simplest solid state model; a good place to read about it is in Feynman’s “Freshmanlectures.”
Fourier transforming c
j=/summationtext
keikajc(k) (where ais the spacing between sites), we
immediately find the energy spectrum ε(k)=− 2tcoska(fig. V .5.2). Imposing a periodic
boundary condition on a lattice with Nsites, we have k=(2π/Na)n withnan integer
from−1
2Nto1
2N.A sN→∞ ,kbecomes a continuous rather than a discrete variable. As
usual, the Brillouin zone is defined by −π/a < k ≤π/a .
There is absolutely nothing relativistic about any of this. Indeed, at the bottom of the
spectrum the energy (up to an irrelevant additive constant) goes as ε(k)/similarequal2t1
2(ka)2≡
k2/2meff. The electron disperses nonrelativistically with an effective mass meff.
jj + 1 j − 1
Figure V .5.1
V .5. Peierls Instability | 299
ε(k)
εF
kπ
a_ +π
a
Figure V .5.2
But now let us fill the system with electrons up to some Fermi energy εF(see fig.
V .5.2). Focus on an electron near the Fermi surface and measure its energy from εF
and momentum from +kF. Suppose we are interested in electrons with energy small
compared to εF, that is, E≡ε−εF/lessmuchεF, and momentum small compared to kF, that is,
p≡k−kF/lessmuchkF. These electrons obey a linear energy-momentum dispersion E=vFp
with the Fermi velocity vF=(∂ε/∂k)| k=kF. We will call the field associated with these
electrons ψR, where the subscript indicates that they are “right moving.” It satisfies the
equation of motion (∂/∂t+vF∂/∂x)ψR=0.
Similarly, the electrons with momentum around −kFobey the dispersion E=−vFp.
We will call the field associated with these electrons ψLwithLfor “left moving,” satisfying
(∂/∂t−vF∂/∂x)ψL=0.
Emergence of the Dirac equation
The Lagrangian summarizing all this is simply
L=iψ†
R/parenleftbigg∂
∂t+vF∂
∂x/parenrightbigg
ψR+iψ†
L/parenleftbigg∂
∂t−vF∂
∂x/parenrightbigg
ψL (2)
Introducing a 2-component field ψ=/parenleftBig
ψL
ψR/parenrightBig
,¯ψ≡ψ†γ0≡ψ†σ2, and choosing units so
thatvF=1, we may write Lmore compactly as
L=iψ†/parenleftbigg∂
∂t−σ3∂
∂x/parenrightbigg
ψ=¯ψiγμ∂μψ (3)
withγ0=σ2andγ1=iσ1satisfying the Clifford algebra {γμ,γν}=2gμν.
300 | V . Field Theory and Collective Phenomena
Amazingly enough, the (1+1)-dimensional Dirac Lagrangian emerges in a totally
nonrelativistic situation!
An instability
I will now go on to discuss an important phenomenon known as Peierls’s instability. I willnecessarily have to be a bit sketchy. I don’t have to tell you again that this is not a text onsolid state physics, but in any case you will not find it difficult to fill in the gaps.
Peierls considered a distortion in the lattice, with the ion at site jdisplaced from its
equilibrium position by cos[ q(ja) ]. (Shades of our mattress from chapter I.1!) A lattice
distortion with wave vector q=2k
Fwill connect electrons with momentum kFwith
electrons with momentum −kF. In other words, it connects right moving ones with
left moving electrons, or in our field theoretic language ψRwithψL. Since the right
moving electrons and the left moving electrons on the surface of the Fermi sea have thesame energy (namely ε
F, duh!) we have the always interesting situation of degenerate
perturbation theory:/parenleftBig
εF 0
0εF/parenrightBig
+/parenleftBig
0δ
δ0/parenrightBig
with eigenvalues εF±δ. A gap opens at the surface
of the Fermi sea. Here δrepresents the perturbation. Thus, Peierls concluded that the
spectrum changes drastically and the system is unstable under a perturbation with wavevector 2 k
F.
A particularly interesting situation occurs when the system is half filled with electrons
(so that the density is one electron per site—recall that electrons have up and down spin).In other words, k
F=π/2aand thus 2 kF=π/a . A lattice distortion of the form shown in
figure V .5.3 has precisely this wave vector. Peierls showed that a half-filled system wouldwant to distort the lattice in this way, doubling the unit cell. It is instructive to see how thisphysical phenomenon emerges in a field theoretic formulation.
Denote the displacement of the ion at site jbyd
j. In the continuum limit, we should be
able to replace djby a scalar field. Show that a perturbation connecting ψRandψLcouples
to¯ψψ and¯ψγ5ψ, and that a linear combination ¯ψψ and¯ψγ5ψcan always be rotated to
¯ψψ by a chiral transformation (see exercise V .5.1.) Thus, we extend (3) to
L=¯ψiγμ∂μψ+1
2[(∂tϕ)2−v2(∂xϕ)2]−1
2μ2ϕ2+gϕ¯ψψ+... (4)
Remember that you worked out the effective potential Veff(ϕ) of this (1+1)-dimensional
field theory in exercise IV .3.2: Veff(ϕ) goes as ϕ2logϕ2for small ϕ, which overwhelms the
1
2μ2ϕ2term. Thus, the symmetry ϕ→−ϕis dynamically broken. The field ϕacquires a
vacuum expectation value and ψbecomes massive. In other words, the electron spectrum
develops a gap.
Figure V .5.3
V .5. Peierls Instability | 301
Exercise
V .5.1 Parallel to the discussion in chapter II.1 you can see easily that the space of 2 by 2 matrices is spanned by
the four matrices I,γμ, andγ5≡γ0γ1=σ3. (Note the peculiar but standard notation of γ5.)Convince
yourself that1
2(I±γ5)projects out right- and left handed fields just as in (3 +1)-dimensional spacetime.
Show that in the bilinear ¯ψγμψleft handed fields are connected to left handed fields and right handed
to right handed and that in the scalar ¯ψψ and the pseudoscalar ¯ψγ5ψright handed is connected to
left handed and vice versa. Finally, note that under the transformation ψ→eiθγ5ψthe scalar and the
pseudoscalar rotate into each other. Check that this transformation leaves the massless Dirac Lagrangian(3) invariant.
V.6 Solitons
Breaking the shackles of Feynman diagrams
When I teach quantum field theory I like to tell the students that by the mid-1970s field
theorists were breaking the shackles of Feynman diagrams. A bit melodramatic, yes,but by that time Feynman diagrams, because of their spectacular successes in quantumelectrodynamics, were dominating the thinking of many field theorists, perhaps to excess.As a student I was even told that Feynman diagrams define quantum field theory, thatquantum fields were merely the “slices of venison”
1used to derive the Feynman rules, and
should be discarded once the rules were obtained. The prevailing view was that it barelymade sense to write down ϕ(x) . This view was forever shattered with the discovery of
topological solitons, as we will now discuss.
Small oscillations versus lumps
Consider once again our favorite toy model L=1
2(∂ϕ)2−V( ϕ) with the infamous double-
well potential V( ϕ)=(λ/4)(ϕ2−v2)2in(1+1)-dimensional spacetime. In chapter IV .1
we learned that of the two vacua ϕ=±vwe are to pick one and study small oscillations
around it. So, pick one and write ϕ=v+χ, expand Linχ, and study the dynamics of
theχmeson with mass μ=(λv2)1
2. Physics then consists of suitably quantized waves
oscillating about the vacuum v.
But that is not the whole story. We can also have a time independent field configuration
withϕ(x) (in this and the next chapter xwill denote only space unless it is clearly meant
to be otherwise from the context) taking on the value −v asx→− ∞ and+v asx→+ ∞ ,
and changing from −v to+v around some point x0over some length scale las shown in
1Gell-Mann used to speak about how pheasant meat is cooked in France between two slices of venison which
are then discarded. He forcefully advocated a program to extract and study the algebraic structure of quantum
field theories which are then discarded.
V .6. Solitons | 303
(a)
(b)+v
−vϕ
x x0
ε
x x0
Figure V .6.1
figure V .6.1a. [Note that if we consider the Euclidean version of the field theory, identify
the time coordinate as the ycoordinate, and think of ϕ(x ,y)as the magnetization (as in
chapter V .3), then the configuration here describes a “domain wall” in a 2-dimensionalmagnetic system.]
Think about the energy per unit length
ε(x)=1
2/parenleftbiggdϕ
dx/parenrightbigg2
+λ
4(ϕ2−v2)2(1)
for this configuration, which I plot in figure V .6.1b. Far away from x0we are in one of the
two vacua and there is no energy density. Near x0, the two terms in ε(x) both contribute to
the energy or mass M=/integraltext
dx ε(x) : the “spatial variation” (in a slight abuse of terminology
often called the “kinetic energy”) term/integraltext
dx1
2(dϕ/dx)2∼l(v/l)2∼v2/l, and the “potential
energy” term/integraltext
dxλ(ϕ2−v2)2∼lλv4. T o minimize the total energy the spatial variation
term wants lto be large, while the potential term wants lto be small. The competition
dM/dl =0 gives v2/l∼lλv4, thus fixing l∼(λv2)−1
2∼1/μ. The mass comes out to be
∼μv2∼μ(μ2/λ).
We have a lump of energy spread over a region of length lof the order of the Compton
wavelength of the χmeson. By translation invariance, the center of the lump x0can be
anywhere. Furthermore, since the theory is Lorentz invariant, we can always boost to
304 | V . Field Theory and Collective Phenomena
send the lump moving at any velocity we like. Recalling a famous retort in the annals
of American politics (“It walks like a duck, quacks like a duck, so Mr. Senator, why don’tyou want to call it a duck?”) we have here a particle, known as a kink or a soliton, withmass ∼μ(μ
2/λ) and size ∼l. Perhaps because of the way the soliton was discovered,
many physicists think of it as a big lumbering object, but as we have seen, the size of asoliton l∼1/μcan be made as small as we like by increasing μ. So a soliton could look like
a point particle. We will come back to this point in chapter VI.3 when we discuss duality.
T opological stability
While the kink and the meson are the same size, for small λthe kink is much more
massive than the meson. Nevertheless, the kink cannot decay into mesons because it costsan infinite amount of energy to undo the kink [by “lifting” ϕ(x) over the potential energy
barrier to change it from +v to−v forxfrom some point >∼x
0to+∞, for example]. The
kink is said to be topologically stable.
The stability is formally guaranteed by the conserved current
Jμ=1
2vεμν∂νϕ (2)
with the charge
Q=/integraldisplay+∞
−∞dxJ0(x)=1
2v[ϕ(+∞)−ϕ(−∞)]
Mesons, which are small localized packets of oscillations in the field clearly have Q=0,
while the kink has Q=1. Thus, the kink cannot decay into a bunch of mesons. Incidentally,
the charge density J0=(1/2v)(dϕ/dx) is concentrated at x0where ϕchanges most rapidly,
as you would expect.
Note that ∂μJμ=0 follows immediately from the antisymmetric symbol εμνand does
not depend on the equation of motion. The current Jμis known as a “topological current.”
Its existence does not follow from Noether’s theorem (chapter I.10) but from topology.
Our discussion also makes clear the existence of an antikink with Q=− 1 and described
by a configuration with ϕ(−∞)=+vandϕ(+∞)=−v. The name is justified by consider-
ing the configuration pictured in figure V .6.2 containing a kink and an antikink far apart.As the kink and the antikink move closer to each other, they clearly can annihilate intomesons, since the configuration shown in figure V .6.2 and the vacuum configuration withϕ(x)=+veverywhere are separated by a finite amount of energy.
A nonperturbative phenomenon
That the mass of the kink comes out inversely proportional to the coupling λis a clear sign
that field theorists could have done perturbation theory in λtill they were blue in the face
without ever discovering the kink. Feynman diagrams could not have told us about it.
V .6. Solitons | 305
Antikink Kink
Figure V .6.2
You can calculate the mass of a kink by minimizing
M=/integraldisplay
dx/bracketleftBigg
1
2/parenleftbiggdϕ
dx/parenrightbigg2
+λ
4/parenleftBig
ϕ2−v2/parenrightBig2/bracketrightBigg
=/parenleftbiggμ2
λ/parenrightbigg
μ/integraldisplay
dy/bracketleftBigg
1
2/parenleftbiggdf
dy/parenrightbigg2
+1
4/parenleftBig
f2−1/parenrightBig2/bracketrightBigg
where in the last step we performed the obvious scaling ϕ(x)→vf (y) andy=μx. This
scaling argument immediately showed that the mass of the kink M=a(μ2/λ)μ with
aa pure number: The heuristic estimate of the mass proved to be highly trustworthy.
The actual function ϕ(x) and hence acan be computed straightforwardly with standard
variational methods.
Bogomol’nyi inequality
More cleverly, observe that the energy density (1) is the sum of two squares. Usinga
2+b2≥2|ab|we obtain
M≥/integraldisplay
dx/parenleftbiggλ
2/parenrightbigg1
2/vextendsingle/vextendsingle/vextendsingle/vextendsingle/parenleftbiggdϕ
dx/parenrightbigg/parenleftBig
ϕ2−v2/parenrightBig/vextendsingle/vextendsingle/vextendsingle/vextendsingle≥/parenleftbiggλ
2/parenrightbigg1
2/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/bracketleftbigg1
3ϕ3−v2ϕ2/bracketrightbigg+∞
−∞/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle=/vextendsingle/vextendsingle/vextendsingle/vextendsingle4
3√
2μ/parenleftbiggμ2
λ/parenrightbigg
Q/vextendsingle/vextendsingle/vextendsingle/vextendsingle
We have the elegant result
M≥|Q| (3)
with mass Mmeasured in units of (4/3√
2)μ(μ2/λ). This is an example of a Bogomol’nyi
inequality, which plays an important role in string theory.
Exercises
V .6.1 Show that if ϕ(x) is a solution of the equation of motion, then so is ϕ[(x−vt)/√
1−v2].
V .6.2 Discuss the solitons in the so-called sine-Gordon theory L=1
2(∂ϕ)2−gcos(βϕ) . Find the topological
current. Is the Q=2 soliton stable or not?
V .6.3 Compute the mass of the kink by the brute force method and check the result from the Bogomol’nyi
inequality.
V.7 Vortices, Monopoles, and Instantons
Vortices
The kink is merely the simplest example of a large class of topological objects in quantum
field theory.
Consider the theory of a complex scalar field in (2+1)-dimensional spacetime L=
∂ϕ†∂ϕ−λ(ϕ†ϕ−v2)2with the now familiar Mexican hat potential. With some minor
changes in notation, this is the theory we used to describe interacting bosons and su-perfluids. (We choose to study the relativistic rather than the nonrelativistic version but asyou will see the issue does not enter for the questions I want to discuss here.)
Are there solitons, that is, objects like the kink, in this theory?Given some time-independent configuration ϕ(x) let us look at its mass or energy
M=/integraldisplay
d2x[∂iϕ†∂iϕ+λ(ϕ†ϕ−v2)2]. (1)
The integrand is a sum of two squares, each of which must give a finite contribution. In
particular, for the contribution of the second term to be finite the magnitude of ϕmust
approach vat spatial infinity.
This finite energy requirement does not fix the phase of ϕhowever. Using polar coordi-
nates(r,θ)we will consider the Ansatz ϕ−→
r→∞veiθ. Writing ϕ=ϕ1+iϕ2, we see that the
vector (ϕ1,ϕ2)=v(cosθ, sinθ)points radially outward at infinity. Recall the definition of
the current Ji=i(∂iϕ†ϕ−ϕ†∂iϕ)in a bosonic fluid given in chapter III.5. The flow whirls
about at spatial infinity, and thus this configuration is known as the vortex.
By explicit differentiation or dimensional analysis, we have ∂iϕ∼v(1/r) asr→∞ .N o w
look at the first term in M. Oops, the energy diverges logarithmically as v2/integraltext
d2x(1/r2).
Is there a way out? Not unless we change the theory.
V .7. Monopoles and Instantons | 307
Vortex into flux tubes
Suppose we gauge the theory by replacing ∂iϕbyDiϕ=∂iϕ−ieAiϕ. Now we can achieve
finite energy by requiring that the two terms in Diϕknock each other out so that Diϕ−→r→∞0
faster than 1 /r. In other words, Ai−→
r→∞−(i/e)(1 /|ϕ|2)ϕ†∂iϕ=(1/e)∂iθ. Immediately, we
have
Flux≡/integraldisplay
d2xF12=/contintegraldisplay
CdxiAi=2π
e(2)
where Cis an infinitely large circle at spatial infinity and we have used Stokes’ theorem.
Thus, in a gauged U(1)theory the vortex carries a magnetic flux inversely proportional to
the charge. When I say magnetic, I am presuming that Arepresents the electromagnetic
gauge potential. The vortex discussed here appears as a flux tube in so-called type IIsuperconductors. It is worth remarking that this fundamental unit of flux (2) is normallywritten in the condensed matter physics literature in unnatural units as
/Phi10=hc
e(3)
very pleasingly uniting three fundamental constants of Nature.
Homotopy groups
Since spatial infinity in 2 dimensional space is topologically a unit circle S1and since
the field configuration with |ϕ|=v also forms a circle S1, this boundary condition can
be characterized as a map S1→S1. Since this map cannot be smoothly deformed into the
trivial map in which S1is mapped onto a point in S1, the corresponding field configuration
is indeed topologically stable. (Think of wrapping a loop of string around a ring.)
Mathematically, maps of Sninto a manifold Mare classified by the homotopy group
/Pi1n(M), which counts the number of topologically inequivalent maps. You can look up the
homotopy groups for various manifolds in tables.1In particular, for n≥1,/Pi1n(Sn)=Z,
where Zis the mathematical notation for the set of all integers. The simplest example
/Pi11(S1)=Zis proved almost immediately by exhibiting the maps ϕ−→r→∞veimθ, with m
any integer (positive or negative), using the context and notation of our discussion for
convenience. Clearly, this map wraps one circle around the other mtimes.
The language of homotopy groups is not just to impress people, but gives us a unifying
language to discuss topological solitons. Indeed, looking back you can now see that thekink is a physical manifestation of /Pi1
0(S0)=Z2, where Z2denotes the multiplicative group
consisting of {+1,−1}(since the 0-dimensional sphere S0={ + 1,−1}consists of just two
points and is topologically equivalent to the spatial infinity in 1-dimensional space).
1See tables 6.V and 6.VI, S. Iyanaga and Y . Kawada, eds., Encyclopedic Dictionary of Mathematics , p. 1415.
308 | V . Field Theory and Collective Phenomena
Hedgehogs and monopoles
If you absorbed all this, you are ready to move up to (3 +1)-dimensional spacetime. Spatial
infinity is now topologically S2. By now you realize that if the scalar field lives on the
manifold M, then we have at infinity the map of S2→M. The simplest choice is thus to
takeS2forM. Hence, we are led to scalar fields ϕa(a=1, 2, 3)transforming as a vector /vectorϕ
under an internal symmetry group O(3)and governed by L=1
2∂/vectorϕ.∂/vectorϕ−V(/vectorϕ./vectorϕ). (There
should be no confusion in using the arrow to indicate a vector in the internal symmetrygroup.)
Let us choose V=λ(/vectorϕ
2−v2)2. The story unfolds much as the story of the vortex. The
requirement that the mass of a time independent configuration
M=/integraldisplay
d3x[1
2(∂/vectorϕ)2+λ(/vectorϕ2−v2)2] (4)
be finite forces |/vectorϕ|=v at spatial infinity so that /vectorϕ(r=∞)indeed lives on S2.
The identity map S2→S2indicates that we should consider a configuration such that
ϕa−→
r→∞vxa
r(5)
This equation looks a bit strange at first sight since it mixes the index of the internal
symmetry group with the index of the spatial coordinates (but in fact we have alreadyencountered this phenomenon in the vortex). At spatial infinity, the field /vectorϕis pointing
radially outward, so this configuration is known picturesquely as a hedgehog. Draw apicture if you don’t get it!
As in the vortex story, the requirement that the first term in (4) be finite forces us to
introduce an O(3)gauge potential A
b
μso that we can replace the ordinary derivative ∂iϕa
by the covariant derivative Diϕa=∂iϕa+eεabcAb
iϕc. We can then arrange Diϕato vanish
at infinity. Simple arithmetic shows that with (5) the gauge potential has to go as
Ab
i−→r→∞1
eεbijxj
r2(6)
Imagine yourself in a lab at spatial infinity. Inside a small enough lab, the /vectorϕfield at
different points are all pointing in approximately the same direction. The gauge groupO(3)is broken down to O(2)/similarequalU(1). The experimentalists in this lab observe a massless
gauge field associated with the U(1), which they might as well call the electromagnetic
field (“quacks like a duck”). Indeed, the gauge invariant tensor field
Fμν≡Fa
μνϕa
|ϕ|−εabcϕa(Dμϕ)b(Dνϕ)c
e|ϕ|3(7)
can be identified as the electromagnetic field (see exercise V .7.5).
There is no electric field since the configuration is time independent and Ab
0=0. We
can only have a magnetic field /vectorBwhich you can immediately calculate since you know Ab
i,
but by symmetry we already see that /vectorBcan point only in the radial direction.
This is the fabled magnetic monopole first postulated by Dirac!
V .7. Monopoles and Instantons | 309
The presence of magnetic monopoles in spontaneously broken gauge theory was dis-
covered by ’t Hooft and Polyakov. If you calculate the total magnetic flux coming out of themonopole/integraltext
d/vectorS./vectorB, where as usual d/vectorSdenotes a small surface element at infinity pointing
radially outward, you will find that it is quantized in suitable units, exactly as Dirac hadstated, as it must (recall chapter IV .4).
We can once again write a Bogomol’nyi inequality for the mass of the monopole
M=/integraldisplay
d3x/bracketleftBig
1
4(/vectorFij)2+1
2(Di/vectorϕ)2+V(/vectorϕ)/bracketrightBig
(8)
[/vectorFijtransforms as a vector under O(3); recall (IV .5.17).] Observe that
1
4(/vectorFij)2+1
2(Di/vectorϕ)2=1
4(/vectorFij±εijkDk/vectorϕ)2∓1
2εijk/vectorFij.Dk/vectorϕ
Thus
M≥/integraldisplay
d3x/bracketleftBig
∓1
2εijk/vectorFij.Dk/vectorϕ+V(/vectorϕ)/bracketrightBig
(9)
We next note that
/integraldisplay
d3x1
2εijk/vectorFij.Dk/vectorϕ=/integraldisplay
d3x1
2εijk∂k(/vectorFij./vectorϕ)=v/integraldisplay
d/vectorS./vectorB=4πvg
has an elegant interpretation in terms of the magnetic charge gof the monopole. Further-
more, if we can throw away V(/vectorϕ)while keeping |/vectorϕ|− →r→∞v, then the inequality M≥4πv|g|
is saturated by /vectorFij=±εijkDk/vectorϕ. The solutions of this equation are known as Bogomol’nyi-
Prasad-Sommerfeld or BPS states.
It is not difficult to construct an electrically charged magnetic monopole, known as a
dyon. We simply take Ab
0=(xb/r)f (r) with some suitable function f( r) .
One nice feature of the topological monopole is that its mass comes out to be ∼MW/α∼
137MW(exercise V .7.11), where MWdenotes the mass of the intermediate vector boson of
the weak interaction. We are anticipating chapter VII.2 a bit in that the gauge boson thatbecomes massive by the Anderson-Higgs mechanism of chapter IV .6 may be identifiedwith the intermediate vector boson. This explains naturally why the monopole has not yetbeen discovered.
Instanton
Consider a nonabelian gauge theory, and rotate the path integral to 4-dimensional Eu-
clidean space. We might wish to evaluate Z=/integraltext
DAe−S(A)in the steepest descent approx-
imation, in which case we would have to find the extrema of
S(A)=/integraldisplay
d4x1
2g2trFμνFμν
with finite action. This implies that at infinity |x|=∞ ,Fμνmust vanish faster than 1 /|x|2,
and so the gauge potential Aμmust be a pure gauge: A=gdg†forgan element of the
gauge group [see (IV .5.6)]. Configurations for which this is true are known as instantons.
310 | V . Field Theory and Collective Phenomena
We see that the instanton is yet one more link in the “great chain of being”: kink-vortex-
monopole-instanton.
Choose the gauge group SU( 2)to be definite. In the parametrization g=x4+i/vectorx./vectorσ
we have by definition g†g=1 and det g=1, thus implying x2
4+/vectorx2=1. We learn that
the group manifold of SU( 2)isS3. Thus, in an instanton, the gauge potential at infinity
A−→
|x|→∞gdg†+O(1/|x|2)defines a map S3→S3. Sound familiar? Indeed, you’ve already
seenS0→S0,S1→S1,S2→S2playing a role in field theory.
Recall from chapter IV .5 that tr F2=dtr(AdA +2
3A3). Thus
/integraldisplay
trF2=/integraldisplay
S3tr(AdA +2
3A3)=/integraldisplay
S3tr(AF−1
3A3)=−1
3/integraldisplay
S3tr(gdg†)3(10)
where we used the fact that Fvanishes at infinity. This shows explicitly that/integraltext
trF2
depends only on the homotopy of the map S3→S3defined by gand is thus a topological
quantity. Incidentally,/integraltext
S3tr(gdg†)3is known to mathematicians as the Pontryagin index
(see exercise V .7.12).
I mentioned in chapter IV .7 that the chiral anomaly is not affected by higher-order
quantum fluctuations. You are now in position to give an elegant topological proof of thisfact (exercise V .7.13).
Kosterlitz-Thouless transition
We were a bit hasty in dismissing the vortex in the nongauged theory L=∂ϕ†∂ϕ−λ(ϕ†ϕ−
v2)2in(2+1)-dimensional spacetime. Around a vortex ϕ∼veiθand so it is true, as we
have noted, that the energy of a single vortex diverges logarithmically. But what about avortex paired with an antivortex?
Picture a vortex and an antivortex separated by a distance Rlarge compared to the
distance scales in the theory. Around an antivortex ϕ∼ve
−iθ. The field ϕwinds around
the vortex one way and around the antivortex the other way. Convince yourself by drawinga picture that at spatial infinity ϕdoes not wind at all: It just goes to a fixed value. The
winding one way cancels the winding the other way.
Thus, a configuration consisting of a vortex-antivortex pair does not cost infinite energy.
But it does cost a finite amount of energy: In the region between the vortex and theantivortex ϕis winding around, in fact roughly twice as fast (as you can see by drawing a
picture). A rough estimate of the energy is thus v
2/integraltext
d2x(1/r2)∼v2log(R/a) , where we
integrated over a region of size R, the relevant physical scale in the problem. (T o make
sense of the problem we divide Rby the size aof the vortex.) The vortex and the antivortex
attract each other with a logarithmic potential. In other words, the configuration cannotbe static: The vortex and the antivortex want to get together and annihilate each other ina fiery embrace and release that finite amount of energy v
2log(R/a). (Hence the term
antivortex.)
All of this is at zero temperature, but in condensed matter physics we are interested in
the free energy F=E−TS (withSthe entropy) at some temperature Trather than the
V .7. Monopoles and Instantons | 311
energy E. Appreciating this elementary point, Kosterlitz and Thouless discovered a phase
transition as the temperature is raised. Consider a gas of vortices and antivortices at somenonzero temperature. Because of thermal agitation, the vortices and antivortices movingaround may or may not find each other to annihilate. How high do we have to crank upthe temperature for this to happen?
Let us do a heuristic estimate. Consider a single vortex. Herr Boltzmann tells us that the
entropy is the logarithm of the “number” of ways in which we can put the vortex insidea box of size L(which we will let tend to infinity). Thus, S∼log(L/a) . The entropy Sis
to battle the energy E∼v
2log(L/a) . We see that the free energy F∼(v2−T)log(L/a)
goes to infinity if T<∼v2, which we identify as essentially the critical temperature Tc.
A single vortex cannot exist below Tc. Vortices and antivortices are tightly bound below
Tcbut are liberated above Tc.
Black hole
The discovery in the 1970s of these topological objects that cannot be seen in perturbation
theory came as a shock to the generation of physicists raised on Feynman diagrams andcanonical quantization. People (including yours truly) were taught that the field operatorϕ(x) is a highly singular quantum operator and has no physical meaning as such, and
that quantum field theory is defined perturbatively by Feynman diagrams. Even quiteeminent physicists asked in puzzlement what a statement such as ϕ−→
r→∞veiθwould mean.
Learned discussions that in hindsight are totally irrelevant ensued. As I said in introducing
chapter V .6, I like to refer to this historical process as “field theorists breaking the shacklesof Feynman diagrams.”
It is worth mentioning one argument physicists at that time used to convince themselves
that solitons do exist. After all, the Schwarzschild black hole, defined by the metric g
μν(x)
(see chapter I.11), had been known since 1916. Just what are the components of the metricg
μν(x)? They are fields in exactly the same way that our scalar field ϕ(x) and our gauge
potential Aμ(x)are fields, and in a quantum theory of gravity gμνwould have to be replaced
by a quantum operator just like ϕandAμ. So the objects discovered in the 1970s are
conceptually no different from the black hole known in the 1910s. But in the early 1970s
most particle theorists were not particularly aware of quantum gravity.
Exercises
V .7.1 Explain the relation between the mathematical statement /Pi10(S0)=Z2and the physical result that there
are no kinks with |Q|≥2.
V .7.2 In the vortex, study the length scales characterizing the variation of the fields ϕandA. Estimate the mass
of the vortex.
312 | V . Field Theory and Collective Phenomena
V .7.3 Consider the vortex configuration in which ϕ−→r→∞veiνθ, with νan integer. Calculate the magnetic flux.
Show that the magnetic flux coming out of an antivortex (for which ν=− 1)is opposite to the magnetic
flux coming out of a vortex.
V .7.4 Mathematically, since g(θ)≡eiνθmay be regarded as an element of the group U(1), we can speak of
a map of S1, the circle at spatial infinity, onto the group U(1). Calculate (i/2π)/integraltext
S1gdg†, thus showing
that the winding number is given by this integral of a 1-form.
V .7.5 Show that within a region in which ϕais constant, Fμνas defined in the text is the electromagnetic field
strength. Compute /vectorBfar from the center of a magnetic monopole and show that Dirac quantization
holds.
V .7.6 Display explicitly the map S2→S2, which wraps one sphere around the other twice. Verify that this map
corresponds to a magnetic monopole with magnetic charge 2.
V .7.7 Write down the variational equations that minimize (8).
V .7.8 Find the BPS solution explicitly.
V .7.9 Discuss the dyon solution. Work it out in the BPS limit.
V .7.10 Verify explicitly that the magnetic monopole is rotation invariant in spite of appearances. By this is
meant that all physical gauge invariant quantities such as /vectorBare covariant under rotation. Gauge variant
quantities such as Ab
ican and do vary under rotation. Write down the generators of rotation.
V .7.11 Show that the mass of the magnetic monopole is about 137 MW.
V .7.12 Evaluate n≡−(1/24π2)/integraltext
S3tr(gdg†)3for the map g=ei/vectorθ./vectorσ. [Hint: By symmetry, you need calculate the
integrand only in a small neighborhood of the identity element of the group or equivalently the north
pole of S3. Next, consider g=ei(θ1σ1+θ 2σ2+mθ 3σ3)forman integer and convince yourself that mmeasures
the number of times S3wraps around S3.] Compare with exercise V .7.4 and admire the elegance of
mathematics.
V .7.13 Prove that higher order corrections do not change the chiral anomaly ∂μJμ
5=[1/(4π)2]εμνλσtrFμνFλσ
(I have rescaled A→(1/g)A). [Hint: Integrate over spacetime and show that the left hand side is given
by the number of right-handed fermion quanta minus the number of left-handed fermion quanta, sothat both sides are given by integers.]
Part VI Field Theory and Condensed Matter
This page intentionally left blank
VI.1Fractional Statistics, Chern-Simons T erm, and
T opological Field Theory
Fractional statistics
The existence of bosons and fermions represents one of the most profound features
of quantum physics. When we interchange two identical quantum particles, the wavefunction acquires a factor of either +1o r −1. Leinaas and Myrheim, and later Wilczek
independently, had the insight to recognize that in (2+1)-dimensional spacetime particles
can also obey statistics other than Bose or Fermi statistics, a statistics now known asfractional or anyon statistics. These particles are now known as anyons.
T o interchange two particles, we can move one of them half-way around the other and
then translate both of them appropriately. When you take one anyon half-way aroundanother anyon going anticlockwise, the wave function acquires a factor of e
iθwhere θ
is a real number characteristic of the particle. For θ=0, we have bosons and for θ=π,
fermions. Particles half-way between bosons and fermions, with θ=π/2, are known as
semions.
After Wilczek’s paper came out, a number of distinguished senior physicists were thor-
oughly confused. Thinking in terms of Schr ¨odinger’s wave function, they got into endless
arguments about whether the wave function must be single valued. Indeed, anyon statis-tics provides a striking example of the fact that the path integral formalism is sometimessignificantly more transparent. The concept of anyon statistics can be formulated in termsof wave functions but it requires thinking clearly about the configuration space over whichthe wave function is defined.
Consider two indistinguishable particles at positions x
i
1andxi
2at some initial time
that end up at positions xf
1andxf
2a time Tlater. In the path integral representation
for/angbracketleftxf
1,xf
2|e−iHT|xi
1,xi
2/angbracketrightwe have to sum over all paths. In spacetime, the worldlines of
the two particles braid around each other (see fig. VI.1.1). (We are implicitly assuming
that the particles cannot go through each other, which is the case if there is a hard corerepulsion between them.) Clearly, the paths can be divided into topologically distinctclasses, characterized by an integer nequal to the number of times the worldlines of the
316 | VI. Field Theory and Condensed Matter
t
i
1xx
i
2xf
2xf
1x
Figure VI.1.1
two particles braid around each other. Since the classes cannot be deformed into each
other, the corresponding amplitudes cannot interfere quantum mechanically, and withthe amplitudes in each class we are allowed to associate an additional phase factor e
iαn
beyond the usual factor coming from the action.
The dependence of αnonnis determined by how quantum amplitudes are to be
combined. Suppose one particle goes around the other through an angle /Delta1ϕ1, a history to
which we assign an additional phase factor eif (/Delta1ϕ 1)withfsome as yet unknown function.
Suppose this history is followed by another history in which our particle goes around theother by an additional angle /Delta1ϕ
2. The phase factor eif (/Delta1ϕ 1+/Delta1ϕ 2)we assign to the combined
history clearly has to satisfy the composition law eif (/Delta1ϕ 1+/Delta1ϕ 2)=eif (/Delta1ϕ 1)eif (/Delta1ϕ 2). In other
words, f (/Delta1ϕ) has to be a linear function of its argument.
We conclude that in (2+1)-dimensional spacetime we can associate with the quantum
amplitude corresponding to paths in which one particle goes around the other anti-clockwise through an angle /Delta1ϕ a phase factor e
i(θ/π)/Delta1ϕ, with θan arbitrary real parameter.
Note that when one particle goes around the other clockwise through an angle /Delta1ϕ the
quantum amplitude acquires a phase factor e−iθ
π/Delta1ϕ.
When we interchange two anyons, we have to be careful to specify whether we do it
“anticlockwise” or “clockwise,” producing factors eiθande−iθ, respectively. This indicates
immediately that parity Pand time reversal invariance Tare violated.
Chern-Simons theory
The next important question is whether all this can be incorporated in a local quantum field
theory. The answer was given by Wilczek and Zee, who showed that the notion of fractionalstatistics can result from the effect of coupling to a gauge potential. The significance of
VI.1. T opological Field Theory | 317
a field theoretic formulation is that it demonstrates conclusively that the idea of anyon
statistics is fully compatible with the cherished principles that we hold dear and that gointo the construction of quantum field theory.
Given a Lagrangian L
0with a conserved current jμ, construct the Lagrangian
L=L0+γεμνλaμ∂νaλ+aμjμ(1)
Hereεμνλdenotes the totally antisymmetric symbol in (2+1)-dimensional spacetime and
γis an arbitrary real parameter. Under a gauge transformation aμ→aμ+∂μ/Lambda1, the term
εμνλaμ∂νaλ, known as the Chern-Simons term, changes by εμνλaμ∂νaλ→εμνλaμ∂νaλ+
εμνλ∂μ/Lambda1∂νaλ. The action changes by δS=γ/integraltext
d3xεμνλ∂μ(/Lambda1∂νaλ)and thus, if we are
allowed to drop boundary terms, as we assume to be the case here, the Chern-Simonsaction is gauge invariant. Note incidentally that in the language of differential form youlearned in chapters IV .5 and IV .6 the Chern-Simons term can be written compactly as ada .
Let us solve the equation of motion derived from (1):
2γεμνλ∂νaλ=−jμ(2)
for a particle sitting at rest (so that ji=0). Integrating the μ=0 component of (2), we
obtain
/integraldisplay
d2x(∂ 1a2−∂2a1)=−1
2γ/integraldisplay
d2xj0(3)
Thus, the Chern-Simons term has the effect of endowing the charged particles in the
theory with flux. (Here the term charged particles simply means particles that couple tothe gauge potential a
μ. In this context, when we refer to charge and flux, we are obviously
not referring to the charge and flux associated with the ordinary electromagnetic field. Weare simply borrowing a useful terminology.)
By the Aharonov-Bohm effect (chapter IV .4), when one of our particles moves around
another, the wave function acquires a phase, thus endowing the particles with anyonstatistics with angle θ=1/4γ(see exercise VI.1.5).
Strictly speaking, the term “fractional statistics” is somewhat misleading. First, a trivial
remark: The statistics parameter θdoes not have to be a fraction. Second, statistics is
not directly related to counting how many particles we can put into a state. The statisticsbetween anyons is perhaps better thought of as a long ranged phase interaction betweenthem, mediated by the gauge potential a.
The appearance of ε
μνλin (1) signals the violation of parity Pand time reversal invari-
anceT, something we already know.
Hopf term
An alternative treatment is to integrate out ain (1). As explained in chapter III.4, and as
in any gauge theory, the inverse of the differential operator ε∂is not defined: It has a zero
mode since (εμνλ∂ν)(∂λF(x)) =0 for any smooth function F(x) . Let us choose the Lorenz
318 | VI. Field Theory and Condensed Matter
gauge ∂μaμ=0. Then, using the fundamental identity of field theory (see appendix A) we
obtain the nonlocal Lagrangian
LHopf=1
4γ/parenleftbigg
jμεμνλ∂ν
∂2jλ/parenrightbigg
(4)
known as the Hopf term.
T o determine the statistics parameter θconsider a history in which one particle moves
half-way around another sitting at rest. The current jis then equal to the sum of two
terms describing the two particles. Plugging into (4) we evaluate the quantum phase
eiS=ei/integraltext
d3xLHopfand obtain θ=1/(4γ).
T opological field theory
There is something conceptually new about the pure Chern-Simons theory
S=γ/integraldisplay
Md3xεμνλaμ∂νaλ (5)
It is topological.
Recall from chapter I.11 that a field theory written in flat spacetime can be immediately
promoted to a field theory in curved spacetime by replacing the Minkowski metric ημνby
the Einstein metric gμνand including a factor√−g in the spacetime integration measure.
But in the Chern-Simons theory ημνdoes not appear! Lorentz indices are contracted
with the totally antisymmetric symbol εμνλ. Furthermore, we don’t need the factor√−g,
as I will now show. Recall also from chapter I.11 that a vector field transforms as aμ(x)=
(∂x/primeλ/∂xμ)a/prime
λ(x/prime)and so for three vector fields
εμνλaμ(x)bν(x)cλ(x)=εμνλ∂x/primeσ
∂xμ∂x/primeτ
∂xν∂x/primeρ
∂xλa/prime
σ(x/prime)b/prime
τ(x/prime)c/prime
ρ(x/prime)
=det/parenleftbigg∂x/prime
∂x/parenrightbigg
εστρa/prime
σ(x/prime)b/prime
τ(x/prime)c/prime
ρ(x/prime)
On the other hand, d3x/prime=d3xdet(∂x/prime/∂x) . Observe, then,
d3xεμνλaμ(x)bν(x)cλ(x)=d3x/primeεστρa/prime
σ(x/prime)b/prime
τ(x/prime)c/prime
ρ(x/prime)
which is invariant without the benefit of√−g.
So, the Chern-Simons action in (5) is invariant under general coordinate transforma-
tion—it is already written for curved spacetime. The metric gμνdoes not enter anywhere.
The Chern-Simons theory does not know about clocks and rulers! It only knows about thetopology of spacetime and is rightly known as a topological field theory. In other words,when the integral in (5) is evaluated over a closed manifold Mthe property of the field
theory/integraltext
Dae
iS(a)depends only on the topology of the manifold, and not on whatever
metric we might put on the manifold.
VI.1. T opological Field Theory | 319
Ground state degeneracy
Recall from chapter I.11 the fundamental definition of energy and momentum. The
energy-momentum tensor is defined by the variation of the action with respect to gμν, but
hey, the action here does not depend on gμν. The energy-momentum tensor and hence the
Hamiltonian is identically zero! One way of saying this is that to define the Hamiltonianwe need clocks and rulers.
What does it mean for a quantum system to have a Hamiltonian H=0? Well, when we
took a course on quantum mechanics, if the professor assigned an exam problem to findthe spectrum of the Hamiltonian 0, we could do it easily! All states have energy E=0. We
are ready to hand it in.
But the nontrivial problem is to find how many states there are. This number is known
as the ground state degeneracy and depends only on the topology of the manifold M.
Massive Dirac fermions and the Chern-Simons term
Consider a gauge potential aμcoupled to a massive Dirac fermion in (2+1)-dimensional
spacetime: L=¯ψ(i/negationslash∂+ /negationslasha−m)ψ . You did an exercise way back in chapter II.1 discovering
the rather surprising phenomenon that in (2+1)-dimensional spacetime the Dirac mass
term violates PandT. (What? You didn’t do it? You have to go back.) Thus, we would expect
to generate the PandTviolating the Chern-Simons term εμνλaμ∂λaνif we integrate out the
fermion to get the term tr log (i/negationslash∂+ /negationslasha−m)in the effective action, along the lines discussed
in chapter IV .3.
In one-loop order we have the vacuum polarization diagram (diagrammatically exactly
the same as in chapter III.7 but in a spacetime with one less dimension) with a Feynmanintegral proportional to
/integraldisplayd3p
(2π)3tr/parenleftbigg
γν 1
/negationslashp+ /negationslashq−mγμ 1
/negationslashp−m/parenrightbigg
(6)
As we will see, the change 4 →3 makes all the difference in the world. I leave it to you
to evaluate (6) in detail (exercise VI.1.7) but let me point out the salient features here.
Since the ∂λin the Chern-Simons term corresponds to qλin momentum space, in order
to identify the coefficient of the Chern-Simons term we need only differentiate (6) withrespect to q
λand set q→0:
/integraldisplayd3p
(2π)3tr(γν 1
/negationslashp−mγλ 1
/negationslashp−mγμ 1
/negationslashp−m)
=/integraldisplayd3p
(2π)3tr[γν(/negationslashp+m)γλ(/negationslashp+m)γμ(/negationslashp+m)]
(p2−m2)3(7)
320 | VI. Field Theory and Condensed Matter
I will simply focus on one piece of the integral, the piece coming from the term in the
trace proportional to m3:
εμνλm3/integraldisplayd3p
(p2−m2)3(8)
As I remarked in exercise II.1.12, in (2 +1)-dimensional spacetime the γμ’s are just the
three Pauli matrices and thus tr (γνγλγμ)is proportional to εμνλ: The antisymmetric
symbol appears as we expect from PandTviolation.
By dimensional analysis, we see that the integral in (8) is up to a numerical constant
equal to 1 /m3and so mcancels.
But be careful! The integral depends only on m2and doesn’t know about the sign of m.
The correct answer is proportional to 1 /|m|3, not 1 /m3. Thus, the coefficient of the Chern-
Simons term is equal to m3/|m|3=m/|m|= sign of m, up to a numerical constant. An
instructive example of an important sign! This makes sense since under P(orT)a Dirac
field with mass mis transformed into a Dirac field with mass −m . In a parity-invariant
theory, with a doublet of Dirac fields with masses mand−m a Chern-Simons term should
not be generated.
Exercises
VI.1.1 In a nonrelativistic theory you might think that there are two separate Chern-Simons terms, εijai∂0ajand
εija0∂iaj. Show that gauge invariance forces the two terms to combine into a single Chern-Simons term
εμνλaμ∂νaλ. For the Chern-Simons term, gauge invariance implies Lorentz invariance. In contrast, the
Maxwell term would in general be nonrelativistic, consisting of two terms, f2
0iandf2
ij, with an arbitrary
relative coefficient between them (with fμν=∂μaν−∂νaμas usual).
VI.1.2 By thinking about mass dimensions, convince yourself that the Chern-Simons term dominates the
Maxwell term at long distances. This is one reason that relativistic field theorists find anyon fluids soappealing. As long as they are interested only in long distance physics they can ignore the Maxwellterm and play with a relativistic theory (see exercise VI.1.1). Note that this picks out (2+1)-dimensional
spacetime as special. In (3+1)-dimensional spacetime the generalization of the Chern-Simons term
ε
μνλσfμνfλσhas the same mass dimension as the Maxwell term f2.I n(4 +1)-dimensional space the
termερμνλσaρfμνfλσis less important at long distances than the Maxwell term f2.
VI.1.3 There is a generalization of the Chern-Simons term to higher dimensional spacetime different from
that given in exercise IV .1.2. We can introduce a p-form gauge potential (see chapter IV .4). Write the
generalized Chern-Simons term in (2p+1)-dimensional spacetime and discuss the resulting theory.
VI.1.4 Consider L=γaε∂a −(1/4g2)f2. Calculate the propagator and show that the gauge boson is massive.
Some physicists puzzled by fractional statistics have reasoned that since in the presence of the Maxwellterm the gauge boson is massive and hence short ranged, it can’t possibly generate fractional statistics,which is manifestly an infinite ranged interaction. (No matter how far apart the two particles we areinterchanging are, the wave function still acquires a phase.) The resolution is that the information is
in fact propagated over an infinite range by a q=0 pole associated with a gauge degree of freedom.
This apparent paradox is intimately connected with the puzzlement many physicists felt when they firstheard of the Aharonov-Bohm effect. How can a particle in a region with no magnetic field whatsoeverand arbitrarily far from the magnetic flux know about the existence of the magnetic flux?
VI.1. T opological Field Theory | 321
VI.1.5 Show that θ=1/4γ. There is a somewhat tricky factor1of 2. So if you are off by a factor of 2, don’t despair.
T ry again.
VI.1.6 Find the nonabelian version of the Chern-Simons term ada. [Hint: As in chapter IV .6 it might be easier
to use differential forms.]
VI.1.7 Using the canonical formalism of chapter I.8 show that the Chern-Simons Lagrangian leads to the
Hamiltonian H=0.
VI.1.8 Evaluate (6).
1X.G. Wen and A. Zee, J. de Physique, 50: 1623, 1989.
VI.2 Quantum Hall Fluids
Interplay between two pieces of physics
Over the last decade or so, the study of topological quantum fluids (of which the Hall fluid
is an example) has emerged as an interesting subject. The quantum Hall system consistsof a bunch of electrons moving in a plane in the presence of an external magnetic fieldBperpendicular to the plane. The magnetic field is assumed to be sufficiently strong so
that the electrons all have spin up, say, so they may be treated as spinless fermions. As iswell known, this seemingly innocuous and simple physical situation contains a wealth ofphysics, the elucidation of which has led to two Nobel prizes. This remarkable richnessfollows from the interplay between two basic pieces of physics.
1. Even though the electron is pointlike, it takes up a finite amount of room.Classically, a charged particle in a magnetic field moves in a Larmor circle of radius
rdetermined by evB=mv
2/r. Classically, the radius is not fixed, with more energetic
particles moving in larger circles, but if we quantize the angular momentum mvr to be
h=2π(in units in which /planckover2piis equal to unity) we obtain eBr2∼2π. A quantum electron
takes up an area of order πr2∼2π2/eB .
2. Electrons are fermions and want to stay out of each other’s way.
Not only does each electron insist on taking up a finite amount of room, each has to
have its own room. Thus, the quantum Hall problem may be described as a sort of housingcrisis, or as the problem of assigning office space at an Institute for Theoretical Physics tovisitors who do not want to share offices.
Already at this stage, we would expect that when the number of electrons N
eis just
right to fill out space completely, namely when Neπr2∼Ne(2π2/eB)∼A, the area of the
system, something special happens.
VI.2. Quantum Hall Fluids | 323
Landau levels and the integer Hall effect
These heuristic considerations could be made precise, of course. The textbook problem of
a single spinless electron in a magnetic field
−[(∂x−ieAx)2+(∂y−ieAy)2]ψ=2mEψ
was solved by Landau decades ago. The states occur in degenerate sets with energy
En=/parenleftbig
n+1
2/parenrightbigeB
m,n=0, 1, 2, ... , known as the nth Landau level. Each Landau level has
degeneracy BA/ 2π, where Ais the area of the system, reflecting the fact that the Larmor
circles may be placed anywhere. Note that one Landau level is separated from the next bya finite amount of energy (eB/m) .
Imagine putting in noninteracting electrons one by one. By the Pauli exclusion principle,
each succeeding electron we put in has to go into a different state in the Landau level. Sinceeach Landau level can hold BA/ 2πelectrons it is natural (see exercise VI.2.1) to define a
filling factor ν≡N
e/(BA/2 π). When νis equal to an integer, the first νLandau levels are
filled. If we want to put in one more electron, it would have to go into the (ν+1)st Landau
level, costing us more energy than what we spent for the preceding electron.
Thus, for νequal to an integer the Hall fluid is incompressible. Any attempt to compress
it lessens the degeneracy of the Landau levels (the effective area Adecreases and so the
degeneracy BA/ 2πdecreases) and forces some of the electrons to the next level, costing
us lots of energy.
An electric field Eyimposed on the Hall fluid in the ydirection produces a current
Jx=σxyEyin the xdirection with σxy=ν(in units of e2/h) . This is easily understood in
terms of the Lorentz force law obeyed by electrons in the presence of a magnetic field. Thesurprising experimental discovery was that the Hall conductance σ
xywhen plotted against
Bgoes through a series of plateaus, which you might have heard about. T o understand
these plateaus we would have to discuss the effect of impurities. I will touch upon thefascinating subject of impurities and disorder in chapter VI.8.
So, the integer quantum Hall effect is relatively easy to understand.
Fractional Hall effect
After the integer Hall effect, the experimental discovery of the fractional Hall effect,
namely that the Hall fluid is also incompressible for filling factor νequal to simple odd-
denominator fractions such as1
3and1
5, took theorists completely by surprise. For ν=1
3,
only one-third of the states in the first Landau level are filled. It would seem that throwing
in a few more electrons would not have that much effect on the system. Why should theν=
1
3Hall fluid be incompressible?
Interaction between electrons turns out to be crucial. The point is that saying the first
Landau level is one-third filled with noninteracting spinless electrons does not define aunique many-body state: there is an enormous degeneracy since each of the electrons can
324 | VI. Field Theory and Condensed Matter
go into any of the BA/ 2πstates available subject only to Pauli exclusion. But as soon as we
turn on a repulsive interaction between the electrons, a presumably unique ground stateis picked out within the vast space of degenerate states. Wen has described the fractionalHall state as an intricate dance of electrons: Not only does each electron occupy a finiteamount of room on the dance floor, but due to the mutual repulsion, it has to be carefulnot to bump into another electron. The dance has to be carefully choreographed, possibleonly for certain special values of ν.
Impurities also play an essential role, but we will postpone the discussion of impurities
to chapter VI.6.
In trying to understand the fractional Hall effect, we have an important clue. You will
remember from chapter V .7 that the fundamental unit of flux is given by 2 π, and thus the
number of flux quanta penetrating the plane is equal to N
φ=BA/ 2π. Thus, the puzzle is
that something special happens when the number of flux quanta per electron Nφ/Ne=ν−1
is an odd integer.
I arranged the chapters so that what you learned in the previous chapter is relevant to
solving the puzzle. Suppose that ν−1flux quanta are somehow bound to each electron.
When we interchange two such bound systems there is an additional Aharonov-Bohmphase in addition to the ( −1)from the Fermi statistics of the electrons. For ν
−1odd these
bound systems effectively obey Bose statistics and can be described by a complex scalarfieldϕ. The condensation of ϕturns out to be responsible for the physics of the quantum
Hall fluid.
Effective field theory of the Hall fluid
We would like to derive an effective field theory of the quantum Hall fluid, first obtainedby Kivelson, Hansson, and Zhang. There are two alternative derivations, a long way and ashort way.
In the long way, we start with the Lagrangian describing spinless electrons in a magnetic
field in the second quantized formalism (we will absorb the electric charge eintoA
μ),
L=ψ+i(∂ 0−iA 0)ψ+1
2mψ†(∂i−iAi)2ψ+V( ψ†ψ) (1)
and massage it into the form we want. In the previous chapter, we learned that by intro-
ducing a Chern-Simons gauge field we can transform ψinto a scalar field. We then invoke
duality, which we will learn about in the next chapter, to represent the phase degree offreedom of the scalar field as a gauge field. After a number of steps, we will discover thatthe effective theory of the Hall fluid turns out to be a Chern-Simons theory.
Instead, I will follow the short way. We will argue by the “what else can it be” method
or, to put it more elegantly, by invoking general principles.
Let us start by listing what we know about the Hall system.
1. We live in (2+1)-dimensional spacetime (because the electrons are restricted to a plane.)
2. The electromagnetic current Jμis conserved: ∂μJμ=0.
VI.2. Quantum Hall Fluids | 325
These two statements are certainly indisputable; when combined they tell us that the
current can be written as the curl of a vector potential
Jμ=1
2π/epsilon1μνλ∂νaλ (2)
The factor of 1 /(2π)defines the normalization of aμ. We learned in school that in 3-
dimensional spacetime, if the divergence of something is zero, then that something isthe curl of something else. That is precisely what (2) says. The only sophistication here isthat what we learned in school works in Minkowskian space as well as Euclidean space—itis just a matter of a few signs here and there.
The gauge potential comes looking for us
Observe that when we transform aμbyaμ→aμ−∂μ/Lambda1, the current is unchanged. In other
words, aμis a gauge potential.
We did not go looking for a gauge potential; the gauge potential came looking for us!
There is no place to hide. The existence of a gauge potential follows from completelygeneral considerations.
3. We want to describe the system field theoretically by an effective local Lagrangian.4. We are only interested in the physics at long distance and large time, that is, at small
wave number and low frequency.
Indeed, a field theoretic description of a physical system may be regarded as a means of
organizing various aspects of the relevant physics in a systematic way according to theirrelative importance at long distances and according to symmetries. We classify terms in afield theoretic Lagrangian according to powers of derivatives, powers of the fields, and soforth. A general scheme for classifying terms is according to their mass dimensions, asexplained in chapter III.2. The gauge potential a
μhas dimension 1, as is always the case for
any gauge potential coupled to matter fields according to the gauge principle, and thus (2)is consistent with the fact that the current has mass dimension 2 in (2+1)-dimensional
spacetime.
5. Parity and time reversal are broken by the external magnetic field.This last statement is just as indisputable as statements 1 and 2. The experimentalist
produces the magnetic field by driving a current through a coil with the current flowingeither clockwise or anticlockwise.
Given these five general statements we can deduce the form of the effective Lagrangian.Since gauge invariance forbids the dimension-2 term a
μaμin the Lagrangian, the
simplest possible term is in fact the dimension-3 Chern-Simons term /epsilon1μνλaμ∂νaλ. Thus,
the Lagrangian is simply
L=k
4πa/epsilon1∂a+... (3)
where kis a dimensionless parameter to be determined.
326 | VI. Field Theory and Condensed Matter
We have introduced and will use henceforth the compact notation /epsilon1a∂b≡/epsilon1μνλaμ∂νbλ=
/epsilon1b∂a for two vector fields aμandbμ.
The terms indicated by ( ...) in (3) include the dimension-4 Maxwell term (1/g2)(f2
0i−
βf2
ij)and other terms with higher dimensions. (Here βis some constant; see exer-
cise VI.1.1.) The important observation is that these higher dimensional terms are lessimportant at long distances. The long distance physics is determined purely by the Chern-Simons term. In general the coefficient kmay well be zero, in which case the physics is
determined by the short distance terms represented by the (...)in (3). Put differently, a
Hall fluid may be defined as a 2-dimensional electron system for which the coefficient ofthe Chern-Simons term does not vanish, and consequently is such that its long distancephysics is largely independent of the microscopic details that define the system. Indeed,we can classify 2-dimensional electron systems according to whether kis zero or not.
Coupling the system to an “external” or “additional” electromagnetic gauge potential A
μ
and using (2) we obtain (after integrating by parts and dropping a surface term)
L=k
4π/epsilon1μνλaμ∂νaλ−1
2π/epsilon1μνλAμ∂νaλ=k
4π/epsilon1μνλaμ∂νaλ−1
2π/epsilon1μνλaμ∂νAλ (4)
Note that the gauge potential of the magnetic field responsible for the Hall effect should
not be included in Aμ; it is implicitly contained already in the coefficient k.
The notion of quasiparticles or “elementary” excitations is basic to condensed matter
physics. The effects of a many-body interaction may be such that the quasiparticles in thesystem are no longer electrons. Here we define the quasiparticles as the entities that coupleto the gauge potential and thus write
L=k
4πa/epsilon1∂a+aμjμ−1
2π/epsilon1μνλaμ∂νAλ... (5)
Defining ˜jμ≡jμ−(1/2π)/epsilon1μνλ∂νAλand integrating out the gauge field we obtain (see
VI.1.4)
L=π
k˜jμ/parenleftbigg/epsilon1μνλ∂ν
∂2/parenrightbigg
˜jλ (6)
Fractional charge and statistics
We can now simply read off the physics from (6). The Lagrangian contains three types
of terms: AA ,Aj, andjj. TheAA term has the schematic form A(/epsilon1∂/epsilon1∂/epsilon1∂/∂2)A. Using
/epsilon1∂/epsilon1∂∼∂2and canceling between numerator and denominator, we obtain
L=1
4πkA/epsilon1∂A (7)
Varying with respect to Awe determine the electromagnetic current
Jμ
em=1
4πk/epsilon1μνλ∂νAλ (8)
VI.2. Quantum Hall Fluids | 327
We learn from the μ=0 component of this equation that an excess density δnof electrons
is related to a local fluctuation of the magnetic field by δn=(1/2πk)δB ; thus we can identify
the filling factor νas 1/k, and from the μ=icomponents that an electric field produces
a current in the orthogonal direction with σxy=(1/k)=ν.
TheAjterm has the schematic form A(/epsilon1∂/epsilon1∂/∂2)j. Canceling the differential operators,
we find
L=1
kAμjμ(9)
Thus, the quasiparticle carries electric charge 1 /k.
Finally, the quasiparticles interact with each other via
L=π
kjμ/epsilon1μνλ∂
∂2jλ(10)
We simply remove the twiddle sign in (6). Recalling chapter VI.1 we see that quasiparticles
obey fractional statistics with
θ
π=1
k(11)
By now, you may well be wondering that while all this is fine and good, what would
actually tell us that ν−1has to be an odd integer?
We now argue that the electron or hole should appear somewhere in the excitation
spectrum. After all, the theory is supposed to describe a system of electrons and thusfar our rather general Lagrangian does not contain any reference to the electron!
Let us look for the hole (or electron). We note from (9) that a bound object made up of
kquasiparticles would have charge equal to 1. This is perhaps the hole! For this to work,
we see that khas to be an integer. So far so good, but kdoesn’t have to be odd yet.
What is the statistics of this bound object? Let us move one of these bound objects half-
way around another such bound object, thus effectively interchanging them. When onequasiparticle moves around another we pick up a phase given by θ/π=1/kaccording to
(11). But here we have kquasiparticles going around kquasiparticles and so we pick up a
phase
θ
π=1
kk2=k (12)
For the hole to be a fermion we must require θ/π to be an odd integer. This fixes kto be
an odd integer.
Sinceν=1/k, we have here the classic Laughlin odd-denominator Hall fluids with filling
factor ν=1
3,1
5,1
7,.... The famous result that the quasiparticles carry fractional charge and
statistics just pops out [see (9) and (11)].
This is truly dramatic: a bunch of electrons moving around in a plane with a magnetic
field corresponding to ν=1
3, and lo and behold, each electron has fragmented into three
pieces, each piece with charge1
3and fractional statistics1
3!
328 | VI. Field Theory and Condensed Matter
A new kind of order
The goal of condensed matter physics is to understand the various states of matter. States
of matter are characterized by the presence (or absence) of order: a ferromagnet becomesordered below the transition temperature. In the Landau-Ginzburg theory, as we saw inchapter V .3, order is associated with spontaneous symmetry breaking, described naturallywith group theory. Girvin and MacDonald first noted that the order in Hall fluids does notreally fit into the Landau-Ginzburg scheme: We have not broken any obvious symmetry.The topological property of the Hall fluids provides a clue to what is going on. As explainedin the preceding chapter, the ground state degeneracy of a Hall fluid depends on thetopology of the manifold it lives on, a dependence group theory is incapable of accountingfor. Wen has forcefully emphasized that the study of topological order, or more generallyquantum order, may open up a vast new vista on the possible states of matter.
1
Comments and generalization
Let me conclude with several comments that might spur you to explore the wealth ofliterature on the Hall fluid.
1. The appearance of integers implies that our result is robust. A slick argument can
be made based on the remark in the previous chapter that the Chern-Simons term doesnot know about clocks and rulers and hence can’t possibly depend on microphysics suchas the scattering of electrons off impurities which cannot be defined without clocks andrulers. In contrast, the physics that is not part of the topological field theory and describedby ( ...) in (3) would certainly depend on detailed microphysics.
2. If we had followed the long way to derive the effective field theory of the Hall fluid, we
would have seen that the quasiparticle is actually a vortex constructed (as in chapter V .7)out of the scalar field representing the electron. Given that the Hall fluid is incompressible,just about the only excitation you can think of is a vortex with electrons coherently whirlingaround.
3. In the previous chapter we remarked that the Chern-Simons term is gauge invariant
only upon dropping a boundary term. But real Hall fluids in the laboratory live in sampleswith boundaries. So how can (3) be correct? Remarkably, this apparent “defect” of the theoryactually represents one of its virtues! Suppose the theory (3) is defined on a bounded 2-dimensional manifold, a disk for example. Then as first argued by Wen there must bephysical degrees of freedom living on the boundary and represented by an action whosechange under a gauge transformation cancels the change of/integraltext
d
3x(k/ 4π)a/epsilon1 ∂a . Physically,
it is clear that an incompressible fluid would have edge excitations2corresponding to waves
on its boundary.
1X. G. Wen, Quantum Field Theory of Many-Body Systems.
2The existence of edge currents in the integer Hall fluid was first pointed out by Halperin.
VI.2. Quantum Hall Fluids | 329
4. What if we refuse to introduce gauge potentials? Since the current Jμhas dimension
2, the simplest term constructed out of the currents, JμJμ, is already of dimension 4;
indeed, this is just the Maxwell term. There is no way of constructing a dimension 3 localinteraction out of the currents directly. T o lower the dimension we are forced to introducethe inverse of the derivative and write schematically J(1//epsilon1∂)J , which is of course just the
non-local Hopf term. Thus, the question “why gauge field?” that people often ask can beanswered in part by saying that the introduction of gauge fields allows us to avoid dealingwith nonlocal interactions.
5. Experimentalists have constructed double-layered quantum Hall systems with an
infinitesimally small tunneling amplitude for electrons to go from one layer to the other.Assuming that the current J
μ
I(I=1, 2)in each layer is separately conserved, we introduce
two gauge potentials by writing Jμ
I=1
2π/epsilon1μνλ∂νaIλas in (2). We can repeat our general
argument and arrive at the effective Lagrangian
L=/summationdisplay
I,JKIJ
4πaI/epsilon1∂aJ+... (13)
The integer khas been promoted to a matrix K. As an exercise, you can derive the Hall
conductance, the fractional charge, and the statistics of the quasiparticles. You would notbe surprised that everywhere 1 /kappears we now have the matrix inverse K
−1instead.
An interesting question is what happens when Khas a zero eigenvalue. For example, we
could have K=/parenleftBig
11
11/parenrightBig
. Then the low energy dynamics of the gauge potential a−≡a1−a2
is not governed by the Chern-Simons term, but by the Maxwell term in the ( ...)in (13).
We have a linearly dispersing mode and thus a superfluid! This striking prediction3was
verified experimentally.
6. Finally, an amusing remark: In this formalism electron tunneling corresponds to the
nonconservation of the current Jμ
−≡Jμ
1−Jμ
2=(1/2π)/epsilon1μνλ∂νa−λ. The difference N1−N2
of the number of electrons in the two layers is not conserved. But how can ∂μJμ
−/negationslash=0 even
though Jμ
−is the curl of a−λ(as I have indicated explicitly)? Recalling chapter IV .4, you the
astute reader say, aha, magnetic monopoles! T unneling in a double-layered Hall system inEuclidean spacetime can be described as a gas of monopoles and antimonopoles.
4(Think,
why monopoles and antimonopoles?) Note of course that these are not monopoles in theusual electromagnetic gauge potential but in the gauge potential a
−λ.
What we have given in this section is certainly a very slick derivation of the effective
long distance theory of the Hall fluid. Some would say too slick. Let us go back to ourfive general statements or principles. Of these five, four are absolutely indisputable. Infact, the most questionable is the statement that looks the most innocuous to the casualreader, namely statement 3. In general, the effective Lagrangian for a condensed mattersystem would be nonlocal. We are implicitly assuming that the system does not contain amassless field, the exchange of which would lead to a nonlocal interaction.
5Also implicit
3X. G. Wen and A. Zee, Phys. Rev. Lett. 69: 1811, 1992.
4X. G. Wen and A. Zee, Phys. Rev. B47: 2265, 1993.
5A technical remark: Vortices (i.e., quasiparticles) pinned to impurities in the Hall fluid can generate an
interaction nonlocal in time.
330 | VI. Field Theory and Condensed Matter
in (3) is the assumption that the Lagrangian can be expressed completely in terms of the
gauge potential a. A priori, we certainly do not know that there might not be other relevant
degrees of freedom. The point is that as long as these degrees of freedom are not gaplessthey can be safely integrated out.
Exercises
VI.2.1 T o define filling factor precisely, we have to discuss the quantum Hall system on a sphere rather than on
a plane. Put a magnetic monopole of strength G(which according to Dirac can be only a half-integer or
an integer) at the center of a unit sphere. The flux through the sphere is equal to Nφ=2G. Show that the
single electron energy is given by El=(1
2/planckover2piωc)/bracketleftbig
l(l+1)−G2/bracketrightbig
/G with the Landau levels corresponding
tol=G,G+1,G+2, . . . , and that the degeneracy of the lth level is 2 l+1. With LLandau levels filled
with noninteracting electrons ( ν=L)show that Nφ=ν−1Ne−S, where the topological quantity Sis
known as the shift.
VI.2.2 For a challenge, derive the effective field theory for Hall fluids with filling factor ν=m/k withkan
odd integer, such as ν=2
5. [Hint: You have to introduce mgauge potentials aIλand generalize (2) to
Jμ=(1/2π)/epsilon1μνλ∂ν/summationtextm
I=1aIλ. The effective theory turns out to be
L=1
4πm/summationdisplay
I,J=1aIKIJ/epsilon1∂aJ+m/summationdisplay
I=1aIμ˜jIμ+...
with the integer kreplaced by a matrix K. Compare with (13).]
VI.2.3 For the Lagrangian in (13), derive the analogs of (8), (9), and (11).
VI.3 Duality
A far reaching concept
Duality is a profound and far reaching concept1in theoretical physics, with origins in
electromagnetism and statistical mechanics. The emergence of duality in recent years inseveral areas of modern physics, ranging from the quantum Hall fluids to string theory,represents a major development in our understanding of quantum field theory. Here Itouch upon one particular example just to give you a flavor of this vast subject.
My plan is to treat a relativistic theory first, and after you get the hang of the subject, I
will go on to discuss the nonrelativistic theory. It makes sense that some of the interestingphysics of the nonrelativistic theory is absent in the relativistic formulation: A largersymmetry is more constraining. By the same token, the relativistic theory is actually mucheasier to understand if only because of notational simplicity.
Vortices
Couple a scalar field in (2+1)-dimensions to an external electromagnetic gauge potential,
with the electric charge qindicated explicitly for later convenience:
L=1
2|(∂μ−iqAμ)ϕ|2−V( ϕ†ϕ) (1)
We have already studied this theory many times, most recently in chapter V .7 in connection
with vortices. As usual, write ϕ=|ϕ|eiθ. Minimizing the potential Vat|ϕ|=v gives the
ground state field configuration. Setting ϕ=veiθin (1) we obtain
L=1
2v2(∂μθ−qAμ)2(2)
1For a first introduction to duality, I highly recommend J. M. Figueroa-O’Farrill, Electromagnetic Duality for
Children , http://www.maths.ed.ac.uk/~jmf/T eaching/Lectures/EDC.html.
332 | VI. Field Theory and Condensed Matter
which upon absorbing θintoAby a gauge transformation we recognize as the Meissner
Lagrangian. For later convenience we also introduce the alternative form
L=−1
2v2ξ2
μ+ξμ(∂μθ−qAμ) (3)
We recover (2) upon eliminating the auxiliary field ξμ(see appendix A and chapter III.5).
In chapter V .7 we learned that the excitation spectrum includes vortices and anti-
vortices, located where |ϕ|vanishes. If, around the zero of |ϕ|,θchanges by 2 π, we have
a vortex. Around an antivortex, /Delta1θ=− 2π. Recall that around a vortex sitting at rest, the
electromagnetic gauge potential has to go as
qAi→∂iθ (4)
at spatial infinity in order for the energy of the vortex to be finite, as we can see from (2).
The magnetic flux
/integraldisplay
d2xεij∂iAj=/contintegraldisplay
d/vectorx./vectorA=/Delta1θvortex
q=2π
q(5)
is quantized in units of 2 π/q .
Let us pause to think physically for a minute. On a distance scale large compared to the
size of the vortex, vortices and antivortices appear as points. As discussed in chapter V .7,the interaction energy of a vortex and an antivortex separated by a distance Ris given
by simply plugging into (2). Ignoring the probe field A
μ, which we can take to be as
weak as possible, we obtain ∼/integraltextR
adr r(∇θ)2∼log(R/a) where ais some short distance
cutoff. But recall that the Coulomb interaction in 2-dimensional space is logarithmic since
by dimensional analysis/integraltext
d2k(ei/vectork./vectorx/k2)∼log(|/vectorx|/a) (with a−1some ultraviolet cutoff).
Thus, a gas of vortices and antivortices appears as a gas of point “charges” with a Coulombinteraction between them.
Vortex as charge in a dual theory
Duality is often made out by some theorists to be a branch of higher mathematics butin fact it derives from an entirely physical idea. In view of the last paragraph, can we notrewrite the theory so that vortices appear as point “charges” of some as yet unknown gaugefield? In other words, we want a dual theory in which the fundamental field creates andannihilates vortices rather than ϕquanta. We will explain the word “dual” in due time.
Remarkably, the rewriting can be accomplished in just a few simple steps. Proceeding
physically and heuristically, we picture the phase field θas smoothly fluctuating, except
that here and there it winds around 2 π. Write ∂
μθ=∂μθsmooth +∂μθvortex . Plugging into
(3) we write
L=−1
2v2ξ2
μ+ξμ(∂μθsmooth +∂μθvortex−qAμ) (6)
Integrate over θsmooth and obtain the constraint ∂μξμ=0, which can be solved by writing
ξμ=εμνλ∂νaλ (7)
VI.3. Duality | 333
a trick we used earlier in chapter VI.2. As in that chapter, a gauge potential comes looking
for us, since the change aλ→aλ+∂λ/Lambda1does not change ξμ. Plugging into (6), we find
L=−1
4v2f2
μν+εμνλ∂νaλ(∂μθvortex−qAμ) (8)
where fμν=∂μaν−∂νaμ.
Our treatment is heuristic because we ignore the fact that |ϕ|vanishes at the vortices.
Physically, we think of the vortices as almost pointlike so that |ϕ|=v “essentially” every-
where. As mentioned in chapter V .6, by appropriate choice of parameters we can makesolitons, vortices, and so on as small as we like. In other words, we neglect the couplingbetween θ
vortex and|ϕ|. A rigorous treatment would require a proper short distance cutoff
by putting the system on a lattice.2But as long as we capture the essential physics, as we
assuredly will, we will ignore such niceties.
Note for later use that the electromagnetic current Jμ, defined as the coefficient of −Aμ
in (8), is determined in terms of the gauge potential aλto be
Jμ=qεμνλ∂νaλ (9)
Let us integrate the term εμνλ∂νaλ∂μθvortex in (8) by parts to obtain aλελμν∂μ∂νθvortex .
According to Newton and Leibniz, ∂μcommutes with ∂ν, and so apparently we get zero.
But∂μand∂νcommute only when acting on a globally defined function, and, heavens to
Betsy, θvortex is not globally defined since it changes by 2 πwhen we go around a vortex.
In particular, consider a vortex at rest and look at the quantity a0couples to in (8) namely
εij∂i∂jθvortex=/vector∇×(/vector∇θvortex)in the notation of elementary physics. Integrating this over
a region containing the vortex gives/integraltext
d2x/vector∇×(/vector∇θvortex)=/contintegraltext
d/vectorx./vector∇θvortex=2π. Thus, we
recognize (1/2π)εij∂i∂jθvortex as the density of vortices, the time component of some vortex
current jλ
vortex. By Lorentz invariance, jλ
vortex=(1/2π)ελμν∂μ∂νθvortex .
Thus, we can now write (8) as
L=−1
4v2f2
μν+(2π)aμjμ
vortex−Aμ(qεμνλ∂νaλ) (10)
Lo and behold, we have accomplished what we set out to do. We have rewritten the theory
so that the vortex appears as an “electric charge” for the gauge potential aμ. Sometimes
this is called a dual theory, but strictly speaking, it is more accurate to refer to it as the dualrepresentation of the original theory (1).
Let us introduce a complex scalar field /Phi1, which we will refer to as the vortex field,
to create and annihilate the vortices and antivortices. In other words, we “elaborate” thedescription in (10) to
L=−1
4v2f2
μν+1
2|(∂μ−i(2π)aμ)/Phi1|2−W(/Phi1) −Aμ(qεμνλ∂νaλ) (11)
2For example, M. P. A. Fisher, “Mott Insulators, Spin Liquids, and Quantum Disordered Superconductivity,”
cond-mat/9806164, appendix A.
334 | VI. Field Theory and Condensed Matter
The potential W(/Phi1) contains terms such as λ(/Phi1†/Phi1)2describing the short distance inter-
action of two vortices (or a vortex and an antivortex.) In principle, if we master all the shortdistance physics contained in the original theory (1) then these terms are all determinedby the original theory.
Vortex of a vortex
Now we come to the most fascinating aspect of the duality representation and the reasonwhy the word “dual” is used in the first place. The vortex field /Phi1is a complex scalar field,
just like the field ϕwe started with. Thus we can perfectly well form a vortex out of /Phi1,
namely a place where /Phi1vanishes and around which the phase of /Phi1goes through 2 π.
Amusingly, we are forming a vortex of a vortex, so to speak.
So, what is a vortex of a vortex?The duality theorem states that the vortex of a vortex is nothing but the original charge,
described by the field ϕwe started out with! Hence the word duality.
The proof is remarkably simple. The vortex in the theory (11) carries “magnetic flux.”
Referring to (11) we see that 2 πa
i→∂iθat spatial infinity. By exactly the same manipulation
as in (5), we have
2π/integraldisplay
d2xεij∂iaj=2π/contintegraldisplay
d/vectorx./vectora=2π (12)
Note that I put quotation marks around the term “magnetic flux” since as is evident I am
talking about the flux associated with the gauge potential aμand not the flux associated
with the electromagnetic potential Aμ. But remember that from (9) the electromagnetic
current Jμ=qεμνλ∂νaλand in particular J0=qεij∂iaj. Hence, the electric charge (note
no quotation marks) of this vortex of a vortex is equal to/integraltext
d2xJ0=q, precisely the charge
of the original complex scalar field ϕ. This proves the assertion.
Here we have studied vortices, but the same sort of duality also applies to monopoles.
As I remarked in chapter IV .4, duality allows us a glimpse into field theories in the stronglycoupled regime. We learned in chapter V .7 that certain spontaneously broken nonabeliangauge theories in (3+1)-dimensional spacetime contains magnetic monopoles. We can
write a dual theory in terms of the monopole field out of which we can construct mono-poles. The monopole of a monopole turns out be none other than the charged fields of theoriginal gauge theory. This duality was first conjectured many years ago by Olive and Mon-tonen and later shown to be realized in certain supersymmetric gauge theories by Seibergand Witten. The understanding of this duality was a “hot” topic a few years ago as it ledto deep new insight about how certain string theories are dual to each other.
3In contrast,
according to one of my distinguished condensed matter colleagues, the important notion
of duality is still underappreciated in the condensed matter physics community.
3For example, D. I. Olive and P. C. West, eds., Duality and Supersymmetric Theories .
VI.3. Duality | 335
Meissner begets Maxwell and so on
We will close by elaborating slightly on duality in (2+1)-dimensional spacetime and how
it might be relevant to the physics of 2-dimensional materials. Consider a Lagrangian L(a)
quadratic in a vector field aμ. Couple an external electromagnetic gauge potential Aμto
the conserved current εμνλ∂νaλ:
L=L(a)+Aμ(εμνλ∂νaλ) (13)
Let us ask: For various choices of L(a), if we integrate out awhat is the effective
Lagrangian L(A) describing the dynamics of A?
If you have gotten this far in the book, you can easily do the integration. The central
identity of quantum field theory again! Given
L(a)∼aKa (14)
we have
L(A)∼(ε∂A)1
K(ε∂A) ∼A(ε∂1
Kε∂)A (15)
We have three choices for L(a) to which I attach various illustrious names:
L(a)∼a2Meissner
L(a)∼aε∂a Chern-Simons (16)
L(a)∼f2∼a∂2aMaxwell
Since we are after conceptual understanding, I won’t bother to keep track of indices and
irrelevant overall constants. (You can fill them in as an exercise.) For example, givenL(a)=f
μνfμνwithfμν=∂μaν−∂νaμwe can write L(a)∼a∂2aand so K=∂2. Thus,
the effective dynamics of the external electromagnetic gauge potential is given by (15) asL(A)∼A[ε∂(1/∂
2)ε∂]A∼A2, the Meissner Lagrangian! In this “quick and dirty” way of
making a living, we simply set εε∼1 and cancel factors of ∂in the numerator against
those in the denominator. Proceeding in this way, we construct the following table:
Dynamics of a K Effective Lagrangian
L(A)∼A[ε∂(1/K)ε∂ ]ADynamics of the
external probe A
Meissner a21A(ε∂ε∂)A ∼A∂2A Maxwell F2
Chern-Simons aε∂a ε∂A(ε∂1
ε∂ε∂)A∼Aε∂A Chern-Simons
Aε∂A
Maxwell f2∼a∂2a∂2A(ε∂1
∂2ε∂)A∼AA Meissner A2
(17)
336 | VI. Field Theory and Condensed Matter
Meissner begets Maxwell, Chern-Simons begets Chern-Simons, and Maxwell begets
Meissner. I find this beautiful and fundamental result, which represents a form of duality,very striking. Chern-Simons is self-dual: It begets itself.
Going nonrelativistic
It is instructive to compare the nonrelativistic treatment of duality.4Go back to the super-
fluid Lagrangian of chapter V .1:
L=iϕ†∂0ϕ−1
2m∂iϕ†∂iϕ−g2(ϕ†ϕ−¯ρ)2(18)
As before, substitute ϕ≡√ρeiθto obtain
L=−ρ∂0θ−ρ
2m(∂iθ)2−g2(ρ−¯ρ)2+... (19)
which we rewrite as
L=−ξμ∂μθ+m
2ρξ2
i−g2(ρ−¯ρ)2+... (20)
In (19) we have dropped a term ∼(∂iρ1/2)2. In (20) we have defined ξ0≡ρ. Integrating
outξiin (20) we recover (19).
All proceeds as before. Writing θ=θsmooth +θvortex and integrating out θsmooth ,w e
obtain the constraint ∂μξμ=0, solved by writing ξμ=/epsilon1μνλ∂νˆaλ. The hat on ˆaλis for later
convenience. Note that the density
ξ0≡ρ=/epsilon1ij∂iˆaj≡ˆf (21)
is the “magnetic” field strength while
ξi=/epsilon1ij(∂0ˆaj−∂jˆa0)≡/epsilon1ijˆf0j (22)
is the “electric” field strength.
Putting all of this into (20) we have
L=m
2ρˆf2
0i−g2(ˆf−¯ρ)2−2πˆaμjμ
vortex+... (23)
T o “subtract out” the background “magnetic” field ¯ρ, an obviously sensible move is to write
ˆaμ=¯aμ+aμ (24)
where we define the background gauge potential by ¯a0=0,∂0¯aj=0 (no background
“electric” field) and
/epsilon1ij∂i¯aj=¯ρ (25)
4The treatment given here follows essentially that given by M. P. A. Fisher and D. H. Lee.
VI.3. Duality | 337
The Lagrangian (23) then takes on the cleaner form
L=/parenleftbiggm
2¯ρf2
0i−g2f2/parenrightbigg
−2πaμjμ
vortex−2π¯aiji
vortex+... (26)
We have expanded ρ∼¯ρin the first term. As in (10) the first two terms form the Maxwell
Lagrangian, and the ratio of their coefficients determines the speed of propagation
c=/parenleftbigg2g2¯ρ
m/parenrightbigg1/2
(27)
In suitable units in which c=1, we have
L=−m
4¯ρfμνfμν−2πaμjμ
vortex−2π¯aiji
vortex+... (28)
Compare this with (10).
The one thing we missed with our relativistic treatment is the last term in (28), for
the simple reason that we didn’t put in a background. Recall that the term like AiJiin
ordinary electromagnetism means that a moving particle associated with the current Ji
sees a magnetic field /vector∇×/vectorA. Thus, a moving vortex will see a “magnetic field”
/epsilon1ij∂i(¯a+a)j=¯ρ+/epsilon1ij∂iaj (29)
equal to the sum of ¯ρ, the density of the original bosons, and a fluctuating field.
In the Coulomb gauge ∂iai=0 we have (f0i)2=(∂0ai)2+(∂ia0)2, where the cross term
(∂0ai)(∂ia0)effectively vanishes upon integration by parts. Integrating out the Coulomb
fielda0, we obtain
L=−¯ρ
2m(2π)2/integraldisplay/integraldisplay
d2xd2y/bracketleftbigg
j0(/vectorx)log|/vectorx−/vectory|
aj0(/vectory)/bracketrightbigg
+m
2¯ρ(∂0ai)2−g2f2+2π(ai+¯ai)jvortex
i(30)
The vortices repel each other by a logarithmic interaction/integraltext
d2k(ei/vectork./vectorx/k2)∼log(|/vectorx|/a) as
we have known all along.
A self-dual theory
Interestingly, the spatial part f2of the Maxwell Lagrangian comes from the short ranged
repulsion between the original bosons.
If we had taken the bosons to interact by an arbitrary potential V( x) we would have,
instead of the last term in (20),
/integraldisplay/integraldisplay
d2xd2y[ρ(x)−¯ρ]V( x−y)[ρ(y)−¯ρ] (31)
It is easy to see that all the steps go through essentially as before, but now the second term
in (26) becomes
/integraldisplay/integraldisplay
d2xd2yf (x)V (x −y)f(y) (32)
338 | VI. Field Theory and Condensed Matter
Thus, the gauge field propagates according to the dispersion relation
ω2=(2¯ρ/m)V (k) /vectork2(33)
where V( k) is the Fourier transformation of V( x) . In the special case V( x)=g2δ(2)(x) we
recover the linear dispersion given in (27). Indeed, we have a linear dispersion ω∝|/vectork|as
long as V( x) is sufficiently short ranged for V(/vectork=0)to be finite.
An interesting case is when V( x) is logarithmic. Then V( k) goes as 1 /k2and so ω∼
constant: The gauge field aibecomes massive and drops out. The low energy effective
theory consists of a bunch of vortices with a logarithmic interaction between them. Thus,a theory of bosons with a logarithmic repulsion between them is self dual in the low energylimit.
The dance of vortices and antivortices
Having gone through this nonrelativistic discussion of duality, let us reward ourselves byderiving the motion of vortices in a fluid. Let the bulk of the fluid be at rest. According to (28)the vortex behaves like a charged particle in a background magnetic field ¯bproportional to
the mean density of the fluid ¯ρ. Thus, the force acting on a vortex is the usual Lorentz force
/vectorv×/vectorB, and the equation of motion of the vortex in the presence of a force Fis then just
¯ρ/epsilon1ij˙xj=Fi (34)
This is the well-known result that a vortex, when pushed, moves in a direction perpendic-
ular to the force.
Consider two vortices. According to (30) they repel each other by a logarithmic inter-
action. They move perpendicular to the force. Thus, they end up circling each other. Incontrast, consider a vortex and an antivortex, which attract each other. As a result of thisattraction, they both move in the same direction, perpendicular to the straight line joiningthem (see fig. VI.3.1). The vortex and antivortex move along in step, maintaining the dis-tance between them. This in fact accounts for the famous motion of a smoke ring. If wecut a smoke ring through its center and perpendicular to the plane it lies in, we have just
++_
(a) (b)+
Figure VI.3.1
VI.3. Duality | 339
a vortex with an antivortex for each section. Thus, the entire smoke ring moves along in a
direction perpendicular to the plane it lies in.
All of this can be understood by elementary physics, as it should be. The key observation
is simply that vortices and antivortices produce circular flows in the fluid around them, sayclockwise for vortices and anticlockwise for antivortices. Another basic observation is thatif there is a local flow in the fluid, then any object, be it a vortex or an antivortex, caught init would just flow along in the same direction as the local flow. This is a consequence ofGalilean invariance. By drawing a simple picture you can see that this produces the samepattern of motion as discussed above.
VI.4 TheσModels as Effective Field Theories
The Lagrangian as a mnemonic
Our beloved quantum field theory has had two near death experiences. The first started
around the mid-1930s when physical quantities came out infinite. But it roared back tolife in the late 1940s and early 1950s, thanks to the work of the generation that includedFeynman, Schwinger, Dyson, and others. The second occurred toward the late 1950s. Aswe have already discussed, quantum field theory seemed totally incapable of addressingthe strong interaction: The coupling was far too strong for perturbation theory to be of anyuse. Many physicists—known collectively as the S-matrix school—felt that field theory
was irrelevant for studying the strong interaction and advocated a program of trying toderive results from general principles without using field theory. For example, in derivingthe Goldberger-T reiman relation, we could have foregone any mention of field theory andFeynman diagrams.
Eventually, in a reaction against this trend, people realized that if some results could be
obtained from general considerations such as notions of spontaneous symmetry breakingand so forth, any Lagrangian incorporating these general properties had to produce thesame results. At the very least, the Lagrangian provides a mnemonic for any physical resultderived without using quantum field theory. Thus was born the notion of long distanceor low energy effective field theory, which would prove enormously useful in both particleand condensed matter physics (as we have already seen and as we will discuss further inchapter VIII.3).
The strong interaction at low energies
One of the earliest examples is the σmodel of Gell-Mann and L ´evy, which describes the
interaction of nucleons and pions. We now know that the strong interaction has to bedescribed in terms of quarks and gluons. Nevertheless, at long distances, the degrees of
VI.4.σModels | 341
freedom are the two nucleons and the three pions. The proton and the neutron transform as
a spinor ψ≡/parenleftbigp
n/parenrightbig
under the SU( 2)of isospin. Consider the kinetic energy term ¯ψiγ∂ψ =
¯ψLiγ∂ψ L+¯ψRiγ∂ψ R. We note that this term has the larger symmetry SU( 2)L×SU( 2)R,
with the left handed field ψLand the right handed field ψRtransforming as a doublet
under SU( 2)LandSU( 2)R, respectively. [The SU( 2)of isospin is the diagonal subgroup
ofSU( 2)L×SU( 2)R.] We can write ψL∼(1
2,0)andψR∼(0,1
2).
Now we see a problem immediately: The mass term m¯ψψ=m(¯ψLψR+h.c.) is not
allowed since ¯ψLψR∼(1
2,1
2), a 4-dimensional representation of SU( 2)L×SU( 2)R,a
group locally isomorphic to SO( 4).
At this point, lesser physicists would have said, what is the problem, we knew all along
that the strong interaction is invariant only under the SU( 2)of isospin, which we will
write as SU( 2)I. Under SU( 2)Ithe bilinears constructed out of ¯ψLandψRtransform as
1
2×1
2=0+1, the singlet being ¯ψψ and the triplet ¯ψiγ5τaψ. With only SU( 2)Isymmetry,
we can certainly include the mass term ¯ψψ .
T o say it somewhat differently, to fully couple to the four bilinears we can construct out
of¯ψLandψR, namely ¯ψψ and¯ψiγ5τaψ, we need four meson fields transforming as the
vector representation under SO( 4). But only the three pion fields are known. It seems
clear that we only have SU( 2)Isymmetry.
Nevertheless, Gell-Mann and L ´evy boldly insisted on the larger symmetry SU( 2)L×
SU( 2)R/similarequalSO( 4)and simply postulated an additional meson field, which they called σ,
so that (σ,/vectorπ)form the 4-dimensional representation. I leave it to you to verify that
¯ψL(σ+i/vectorτ./vectorπ)ψR+h.c.=¯ψ(σ+i/vectorτ./vectorπγ 5)ψis invariant. Hence, we can write down the
invariant Lagrangian
L=¯ψ[iγ∂+g(σ+i/vectorτ./vectorπγ 5)]ψ+L(σ,/vectorπ) (1)
where the part not involving the nucleons reads
L(σ,/vectorπ)=1
2/parenleftBig
(∂σ)2+(∂/vectorπ)2/parenrightBig
+μ2
2(σ2+/vectorπ2)−λ
4(σ2+/vectorπ2)2(2)
This is known as the linear σmodel.
Theσmodel would have struck most physicists as rather strange at the time it was
introduced: The nucleon does not have a mass and there is an extra meson field. Aha, butyou would recognize (2) as precisely the Lagrangian (IV .1.2) (for N=4)that we studied,
which exhibits spontaneous symmetry breaking. The four scalar fields (ϕ
4,ϕ1,ϕ2,ϕ3)
in (IV .1.2) correspond to (σ,/vectorπ). With no loss of generality, we can choose the vacuum
expectation value of ϕto point in the 4th direction, namely the vacuum in which /angbracketleft0|σ|0/angbracketright=/radicalbig
μ2/λ≡vand/angbracketleft0|/vectorπ|0/angbracketright=0. Expanding σ=v+σ/primewe see immediately that the nucleon
has a mass M=gv. You should not be surprised that the pion comes out massless. The
meson associated with the field σ/prime, which we will call the σmeson, has no reason to be
massless and indeed is not.
Can the all-important parameter vbe related to a measurable quantity? Indeed. From
chapter I.10 you will recall that the axial current is given by Noether’s theorem asJ
a
μ5=¯ψγμγ5(τa/2)ψ+πa∂μσ−σ∂μπa. After σacquires a vacuum expectation value,
342 | VI. Field Theory and Condensed Matter
Ja
μ5contains a term −v∂μπa. This term implies that the matrix element /angbracketleft0|Ja
μ5|πb/angbracketright=ivk μ,
where kdenotes the momentum of the pion, and thus vis proportional to the fdefined in
chapter IV .2. Indeed, we recognize the mass relation M=gvas precisely the Goldberger-
T reiman relation (IV .2.7) with F(0)=1 (see exercise VI.4.4).
The nonlinear σmodel
It was eventually realized that the main purpose in life of the potential in L(σ,/vectorπ)is to
force the vacuum expectation values of the fields to be what they are, so the potential canbe replaced by a constraint σ
2+/vectorπ2=v2. A more physical way of thinking about this point
is by realizing that the σmeson, if it exists at all, must be very broad since it can decay
via the strong interaction into two pions. We might as well force it out of the low energyspectrum by making its mass large. By now, you have learned from chapters IV .1 and V .1that the mass of the σmeson, namely√
2μ, can be taken to infinity while keeping vfixed
by letting μ2andλtend to infinity, keeping their ratio fixed.
We will now focus on L(σ,/vectorπ). Instead of thinking abut L(σ,/vectorπ)=1
2[(∂σ)2+(∂/vectorπ)2]
with the constraint σ2+/vectorπ2=v2, we can simply solve the constraint and plug the solution
σ=√
v2−/vectorπ2into the Lagrangian, thus obtaining what is known as the nonlinear σmodel:
L=1
2/bracketleftbigg
(∂/vectorπ)2+(/vectorπ.∂/vectorπ)2
f2−/vectorπ2/bracketrightbigg
=1
2(∂/vectorπ)2+1
2f2(/vectorπ.∂/vectorπ)2+... (3)
Note that Lcan be written in the form L=(∂πa)Gab(/vectorπ)(∂πb); some people like to think of
Gabas a “metric” in field space. [Incidentally, recall that way back in chapter I.3 we restricted
ourselves to the simplest possible kinetic energy term1
2(∂ϕ)2, rejecting possibilities such
asU(ϕ)(∂ϕ)2. But recall also that in chapter IV .3 we noted that such a term would arise by
quantum fluctuations.]
In accordance with the philosophy that introduced this chapter, any Lagrangian that
captures the correct symmetry properties should describe the same low energy physics.1
This means that anybody, including you, can introduce his or her own parametrization ofthe fields.
The nonlinear σmodel is actually an example of a broad class of field theories whose
Lagrangian has a simple form but with the fields appearing in it subject to some nontrivialconstraint. An example is the theory defined by
L(U)=f2
4tr(∂μU†.∂μU) (4)
withU(x) a matrix-valued field and an element of SU( 2). Indeed, if we write U=e(i/f )/vectorπ./vectorτ
we see that L(U)=1
2(∂/vectorπ)2+(1/2f2)(/vectorπ.∂/vectorπ)2+... , identical to (3) up to the terms
indicated. The /vectorπfield here is related to the one in (3) by a field redefinition.
There is considerably more we can say about the nonlinear σmodels and their applica-
tions in particle and condensed matter physics, but a thorough discussion would take us
1S. Weinberg, Physica 96A: 327, 1979.
VI.4.σModels | 343
far beyond the scope of this book. Instead, I will develop some of their properties in the
exercises and in the next chapter will sketch how they can arise in one class of condensedmatter systems.
Exercises
VI.4.1 Show that the vacuum expectation value of (σ,/vectorπ)can indeed point in any direction without changing the
physics. At first sight, this statement seems strange since, by virtue of its γ5coupling to the nucleon, the
pion is a pseudoscalar field and cannot have a vacuum expectation without breaking parity. But (σ,/vectorπ)
are just Greek letters. Show that by a suitable transformation of the nucleon field parity is conserved, asit should be in the strong interaction.
VI.4.2 Calculate the pion-pion scattering amplitude up to quadratic order in the external momenta, using the
nonlinear σmodel (3). [Hint: For help, see S. Weinberg, Phys. Rev. Lett. 17: 616, 1966.]
VI.4.3 Calculate the pion-pion scattering amplitude up to quadratic order in the external momenta, using the
linear σmodel (2). Don’t forget the Feynman diagram involving σmeson exchange. You should get the
same result as in exercise VI.4.2.
VI.4.4 Show that the mass relation M=gvamounts to the Goldberger-T reiman relation.
VI.5 Ferromagnets and Antiferromagnets
Magnetic moments
In chapters IV .1 and V .3 I discussed how the concept of the Nambu-Goldstone boson
originated as the spin wave in a ferromagnetic or an antiferromagnetic material. A cartoondescription of such materials consists of a regular lattice on each site of which sits alocal magnetic moment, which we denote by a unit vector /vectorn
jwithjlabeling the site.
In a ferromagnetic material the magnetic moments on neighboring sites want to pointin the same direction, while in an antiferromagnetic material the magnetic momentson neighboring sites want to point in opposite directions. In other words, the energy isH=J/summationtext
<ij>/vectorni./vectornj, where iandjlabel neighboring sites. For antiferromagnets J>
0, and for ferromagnets J< 0. I will merely allude to the fully quantum description
formulated in terms of a spin /vectorSjoperator on each site j; the subject lies far beyond the
scope of this text.
In a more microscopic treatment, we would start with a Hamiltonian (such as the
Hubbard Hamiltonian) describing the hopping of electrons and the interaction betweenthem. Within some approximate mean field treatment the classical variable /vectorn
jwould then
emerge as the unit vector pointing in the direction of /angbracketleftc†
j/vectorσcj/angbracketrightwithc†
jandcjthe electron
creation and annihilation operators, respectively. But this is not a text on solid state physics.
First versus second order in time
Here we would like to derive an effective low energy description of the ferromagnet and
antiferromagnet in the spirit of the σmodel description of the preceding chapter. Our
treatment will be significantly longer than the standard discussion given in some fieldtheory texts, but has the slight advantage of being correct.
The somewhat subtle issue is what kinetic energy term we have to add to −H to form the
Lagrangian L. Since for a unit vector /vectornwe have /vectorn.(d/vectorn/dt) =(d(/vectorn./vectorn)/dt) =0, we cannot
VI.5. Magnetic Systems | 345
make do with one time derivative. With two derivatives we can form (d/vectorn/dt) .(d/vectorn/dt)
and so
Lwrong=1
2g2/summationdisplay
j∂/vectornj
∂t.∂/vectornj
∂t−J/summationdisplay
<ij>/vectorni./vectornj (1)
A typical field theory text would then pass to the continuum limit and arrive at the
Lagrangian density
L=1
2g2(∂/vectorn
∂t.∂/vectorn
∂t−c2
s/summationdisplay
l∂/vectorn
∂xl.∂/vectorn
∂xl) (2)
with the constraint [ /vectorn(x ,t)]2=1. This is another example of a nonlinear σmodel. Just as
in the nonlinear σmodel discussed in chapter VI.4, the Lagrangian looks free, but the
nontrivial dynamics comes from the constraint. The constant cs(which is determined in
terms of the microscopic variable J)is the spin wave velocity, as you can see by writing
down the equation of motion (∂2/∂t2)/vectorn−c2
s∇2/vectorn=0.
But you can feel that something is wrong. You learned in a quantum mechanics course
that the dynamics of a spin variable /vectorSis first order in time. Consider the most basic
example of a spin in a constant magnetic field described by H=μ/vectorS./vectorB. Then d/vectorS/dt=
i[H,/vectorS]=μ/vectorB×/vectorS. Besides, you might remember from a solid state physics course that in
a ferromagnet the dispersion relation of the spin wave has the nonrelativistic form ω∝k2
and not the relativistic form ω2∝k2implied by (2).
The resolution of this apparent paradox is based on the Pauli-Hopf identity: Given a unit
vector /vectornwe can always write /vectorn=z†/vectorσz, where z=/parenleftbigz1
z2/parenrightbig
consists of two complex numbers
such that z†z≡z†
1z1+z†
2z2=1. Verify this! (A mathematical aside: Writing z1andz2out
in terms of real numbers we see that this defines the so-called Hopf map S3→S2.)While
we cannot form a term quadratic in /vectornand linear in time derivative, we can write a term
quadratic in the complex doublet zand linear in time derivative. Can you figure it out
before looking at the next line?
The correct version of (1) is
Lcorrect=i/summationdisplay
jz†
j∂zj
∂t+1
2g2/summationdisplay
j∂/vectornj
∂t.∂/vectornj
∂t−J/summationdisplay
<ij>/vectorni./vectornj (3)
The added term is known as the Berry’s phase term and has deep topological meaning.
You should derive the equation of motion using the identity
/integraldisplay
dt δ/parenleftbigg
z†
j∂zj
∂t/parenrightbigg
=1
2i/integraldisplay
dt δ/vectornj./parenleftbigg
/vectornj×∂/vectornj
∂t/parenrightbigg
(4)
Remarkably, although z†
j(∂zj/∂t) cannot be written simply in terms of /vectornj, its variation
can be.
Low energy modes in the ferromagnet and the antiferromagnet
In the ground state of a ferromagnet, the magnetic moments all point in the same
direction, which we can choose to be the z-direction. Expanding the equation of motion in
346 | VI. Field Theory and Condensed Matter
small fluctuations around this ground state /vectornj=ˆez+δ/vectornj(where evidently ˆezdenotes the
appropriate unit vector) and Fourier transforming, we obtain
/parenleftBigg−ω2
g2+h(k) −1
2iω
1
2iω −ω2
g2+h(k)/parenrightBigg/parenleftBiggδnx(k)
δny(k)/parenrightBigg
=0 (5)
linking the two components δnx(k)andδny(k)ofδ/vectorn(k) . The condition /vectornj./vectornj=1 says that
δnz(k)=0. Here ais the lattice spacing and h(k)≡4J[2−cos(kxa)−cos(kya)]/similarequal2Ja2k2
for small k. (I am implicitly working in two spatial dimensions as evidenced by kxandky.)
At low frequency the Berry term iωdominates the naive term ω2/g2, which we can
therefore throw away. Setting the determinant of the matrix equal to zero, we see that weget the correct quadratic dispersion relation ω∝k
2.
The treatment of the antiferromagnet is interestingly different. The so-called N ´eel state1
for an antiferromagnet is defined by /vectornj=(−1)jˆez. Writing /vectornj=(−1)jˆez+δ/vectornj, we obtain
/parenleftBigg−ω2
g2+f( k) −1
2iω
1
2iω −ω2
g2+f( k)/parenrightBigg/parenleftBiggδnx(k)
δny(k+Q)/parenrightBigg
=0 (6)
linking δnx and δny evaluated at different momenta. Here f( k)=4J
[2+cos(kxa)+cos(kya)] andQ=[π/a ,π/a ]. The appearance of Qis due to (−1)j=eiQaj.
(I will let you figure out the somewhat overly compact notation.) The antiferromagneticfactor (−1)
jexplicitly breaks translation invariance and kicks in the momentum Qwhen-
ever it occurs. A similar equation links δny(k) andδnx(k+Q). Solving these equations,
you will find that there is a high frequency branch that we are not interested in and a lowfrequency branch with the linear dispersion ω∝k.
Thus, the low frequency dynamics of the antiferromagnet can be described by the
nonlinear σmodel (2), which when the spin wave velocity is normalized to 1 can be written
in the relativistic form:
L=1
2g2∂μ/vectorn.∂μ/vectorn (7)
Exercises
VI.5.1 Work out the two branches of the spin wave spectrum in the ferromagnetic case, paying particular
attention to the polarization.
VI.5.2 Verify that in the antiferromagnetic case the Berry’s phase term merely changes the spin wave velocity
and does not affect the spectrum qualitatively as in the ferromagnetic case.
1Note that while the N ´eel state describes the lowest energy configuration for a classical antiferromagnet, it
does not describe the ground state of a quantum antiferromagnet. The terms S+
iS−
j+S−
iS+
jin the Hamiltonian
J/summationtext
<ij>/vectorSi./vectorSjflip the spins up and down.
VI.6 Surface Growth and Field Theory
In this chapter I will discuss a topic, rather unusual for a field theory text, taken from non-
equilibrium statistical mechanics, one of the hottest growth fields in theoretical physicsover the last few years. I want to introduce you to yet another area in which field theoreticconcepts are of use.
Imagine atoms being deposited randomly on some surface. This is literally how some
novel materials are grown. The height h(x ,t)of the surface as it grows is governed by the
Kardar-Parisi-Zhang equation
∂h
∂t=ν∇2h+λ
2(∇h)2+η(/vectorx,t) (1)
This equation describes a deceptively simple prototype of nonequilibrium dynamics and
has a remarkably wide range of applicability.
T o understand (1), consider the various terms on the right-hand side. The term ν∇2h
(withν> 0)is easy to understand: Positive in the valleys of hand negative on the peaks,
it tends to smooth out the surface. With only this term the problem would be linear andhence trivial. The nonlinear term (λ/2)(∇h)
2renders the problem highly nontrivial and
interesting; I leave it to you as an exercise to convince yourself of the geometric origin ofthis term. The third term describes the random arrival of atoms, with the random variableη(/vectorx,t)usually assumed to be Gaussian distributed, with zero mean,
1and correlations
/angbracketleftbig
η(/vectorx,t)η(/vectorx/prime,t/prime)/angbracketrightbig
=2σ2δD(/vectorx−/vectorx/prime)δ(t−t/prime) (2)
In other words, the probability distribution for a particular η(/vectorx,t)is given by
P(η)∝e−1
2σ2/integraltext
dDxdt η( /vectorx,t)2
Here /vectorxrepresents coordinates in D-dimensional space. Experimentally, D=2 for the
situation I described, but theoretically we are free to investigate the problem for any D.
1There is no loss of generality here since an additive constant in ηcan be absorbed by the shift h→h+ct.
348 | VI. Field Theory and Condensed Matter
Typically, condensed matter physicists are interested in calculating the correlation be-
tween the height of the surface at two different positions in space and time:
/angbracketleft[h(/vectorx,t)−h(/vectorx/prime,t/prime)]2/angbracketright=|/vectorx−/vectorx/prime|2χf/parenleftbigg|/vectorx−/vectorx/prime|z
|t−t/prime|/parenrightbigg
, (3)
The bracket /angbracketleft.../angbracketrighthere and in (2) denotes averaging over different realizations of the
random variable η(/vectorx,t). On the right-hand side of (3) I have written the dynamic scaling
form typically postulated in condensed matter physics, where χandzare the so-called
roughness and dynamic exponents. The challenge is then to show that the scaling form iscorrect and to calculate χandz. Note that the dynamic exponent z(which in general is not
an integer) tells us, roughly speaking, how many powers of space is worth one power oftime. (For λ=0 we have simple diffusion for which z=2.)Herefdenotes an unknown
function.
I will not go into more technical details. Our interest here is to see how this problem,
which does not even involve quantum mechanics, can be converted into a quantum fieldtheory. Start with
Z≡/integraldisplay
Dh/integraldisplay
Dηe−1
2σ2/integraltext
dDxdt η( /vectorx,t)2
δ/bracketleftbigg∂h
∂t−ν∇2h−λ
2(∇h)2−η(/vectorx,t)/bracketrightbigg
(4)
Integrating over η, we obtain Z=/integraltext
Dhe−S(h)with the action
S(h)=1
2σ2/integraldisplay
dD/vectorxd t/bracketleftbigg∂h
∂t−ν∇2h−λ
2(∇h)2/bracketrightbigg2
(5)
You will recognize that this describes a nonrelativistic field theory of a scalar field h(/vectorx,t).
The physical quantity we are interested in is then given by
/angbracketleft/bracketleftbig
h(/vectorx,t)−h(/vectorx/prime,t/prime)/bracketrightbig2/angbracketright=1
Z/integraldisplay
Dhe−S(h)[h(/vectorx,t)−h(/vectorx/prime,t/prime)]2(6)
Thus, the challenge of determing the roughness and dynamic exponents in statistical
physics is equivalent to the problem of determining the propagator
D(/vectorx,t)≡1
Z/integraldisplay
Dhe−S(h)h(/vectorx,t)h(/vector0, 0)
of the scalar field h.
Incidentally, by scaling t→t/ν andh→/radicalbig
σ2/ν h, we can write the action as
S(h)=1
2/integraldisplay
dD/vectorxd t/bracketleftbigg/parenleftbigg∂
∂t−∇2/parenrightbigg
h−g
2(∇h)2/bracketrightbigg2
(7)
withg2≡λ2σ2/ν3. Expanding the action in powers of has usual
S(h)=1
2/integraldisplay
dD/vectorxd t/braceleftBigg/bracketleftbigg/parenleftbigg∂
∂t−∇2/parenrightbigg
h/bracketrightbigg2
−g(∇h)2/parenleftbigg∂
∂t−∇2/parenrightbigg
h+g2
4(∇h)4/bracerightBigg
,
(8)
we recognize the quadratic term as giving us the rather unusual propagator 1 /(ω2+k4)for
the scalar field h, and the cubic and quartic term as describing the interaction. As always,
VI.6. Surface Growth | 349
h
h(x)
θdcos θ
x
Figure VI.6.1
to calculate the desired physical quantity we evaluate the functional or “path” integral
Z=/integraldisplay
Dhe−S(h)+/integraltext
dDxdtJ(x ,t)h(x ,t)
and then functionally differentiate repeatedly with respect to J.
My intent here is not so much to teach you nonequilibrium statistical mechanics as to
show you that quantum field theory can emerge in a variety of physical situations, includingthose involving only purely classical physics. Note that the “quantum fluctuations” herearise from the random driving term. Evidently, there is a close methodological connectionbetween random dynamics and quantum physics.
Exercises
VI.6.1 An exercise in elementary geometry: Draw a straight line tilted at an angle θwith respect to the horizontal.
The line represents a small segment of the surface at time t. Now draw a number of circles of diameter
dtangent to and on top of this line. Next draw another line tilted at angle θwith respect to the horizontal
and lying on top of the circles, namely tangent to them. This new line represents the segment of thesurface some time later (see fig. VI.6.1). Note that /Delta1h=d/cosθ/similarequald(1+
1
2θ2). Show that this generates
the nonlinear term (λ/2)(∇h)2in the KPZ equation (1). For applications of the KPZ equation, see for
example, T . Halpin–Healy and Y .-C. Zhang, Phys. Rep. 254: 215, 1995; A. L. Barabasi and H. E. Stanley,
Fractal Concepts in Surface Growth .
VI.6.2 Show that the scalar field hhas the propagator 1 /(ω2+k4).
VI.6.3 Field theory can often be cast into apparently rather different forms by a change of variable. Show that
by writing U=e1
2ghwe can change the action (7) to
S=2
g2/integraldisplay
dD/vectorxd t/parenleftbigg
U−1∂
∂tU−U−1∇2U/parenrightbigg2
(9)
a kind of nonlinear σmodel.
VI.7 Disorder: Replicas and Grassmannian Symmetry
Impurities and random potential
An important area in condensed matter physics involves the study of disordered systems,
a subject that has been the focus of a tremendous amount of theoretical work over thelast few decades. Electrons in real materials scatter off the impurities inevitably presentand effectively move in a random potential. In the spirit of this book I will give you a briefintroduction to this fascinating subject, showing how the problem can be mapped into aquantum field theory.
The prototype problem is that of a quantum particle obeying the Schr ¨odinger equa-
tionHψ=[−∇
2+V( x) ]ψ=Eψ , where V( x) is a random potential (representing the
impurities) generated with the Gaussian white noise probability distribution P(V) =
Ne−/integraltext
dDx(1/2g2)V (x)2with the normalization factor Ndetermined by/integraltext
DVP(V) =1. The
parameter gmeasures the strength of the impurities: the larger g, the more disordered the
system. This of course represents an idealization in which interaction between electronsand a number of other physical effects are neglected.
As in statistical mechanics we think of an ensemble of systems each of which is charac-
terized by a particular function V( x) taken from the distribution P(V) . We study the aver-
age or typical properties of the system. In particular, we might want to know the averageddensity of states defined by ρ(E)=/angbracketlefttrδ(E−H)/angbracketright=/angbracketleft/summationtext
iδ(E−Ei)/angbracketright, where the sum runs
over the ith eigenstate of Hwith corresponding eigenvalue Ei. We denote by /angbracketleftO(V) /angbracketright≡/integraltext
DVP( V) O( V) the average of any functional O(V) ofV( x) . Clearly,/integraltextE∗+δE
E∗dE ρ(E)
counts the number of states in the interval from E∗toE∗+δE, an important quantity in,
for example, tunneling experiments.
VI.7. Disorder | 351
Anderson localization
Another important physical question is whether the wave functions at a particular energy
Eextend over the entire system or are localized within a characteristic length scale ξ(E) .
Clearly, this issue determines whether the material is a conductor or an insulator. At firstsight, you might think that we should study
S(x ,y;E)≡/angbracketleftBig/summationdisplay
iδ(E−Ei)ψ∗
i(x)ψi(y)/angbracketrightBig
which might tell us how the wave function at xis correlated with the wave function at
some other point y, butSis unsuitable because ψ∗
i(x)ψ i(y) has a phase that depends on
V. Thus, Swould vanish when averaged over disorder. Instead, the correct quantity to
study is
K(x−y;E)≡/angbracketleftBig/summationdisplay
iδ(E−Ei)ψ∗
i(x)ψi(y)ψ∗
i(y)ψi(x)/angbracketrightBig
sinceψ∗
i(x)ψi(y)ψ∗
i(y)ψi(x)is manifestly positive. Note that upon averaging over all possi-
bleV( x) we recover translation invariance so that Kdoes not depend on xandyseparately,
but only on the separation |x−y|.A s|x −y|→∞ ,i fK(x−y;E)∼e−|x−y|/ξ(E)decreases
exponentially the wave functions around the energy Eare localized over the so-called local-
ization length ξ(E) . On the other hand, if K(x−y;E)decreases as a power law of |x−y|,
the wave functions are said to be extended.
Anderson and his collaborators made the surprising discovery that localization prop-
erties depend on D, the dimension of space, but not on the detailed form of P(V) (an
example of the notion of universality). For D=1 and 2 all wave functions are localized,
regardless of how weak the impurity potential might be. This is a highly nontrivial state-ment since a priori you might think, as eminent physicists did at the time, that whether thewave functions are localized or not depends on the strength of the potential. In contrast, forD=3, the wave functions are extended for Ein the range (−E
c,Ec).A sEapproaches the
energy Ec(known as the mobility edge) from above, the localization length ξ(E) diverges
asξ(E)∼1/(E−Ec)μwith some critical exponent1μ. Anderson received the Nobel Prize
for this work and for other contributions to condensed matter theory.
Physically, localization is due to destructive interference between the quantum waves
scattering off the random potential.
When a magnetic field is turned on perpendicular to the plane of a D=2 electron gas
the situation changes dramatically: An extended wave function appears at E=0. For non-
zeroE, all wave functions are still localized, but with the localization length diverging as
ξ(E)∼1/|E|ν. This accounts for one of the most striking features of the quantum Hall
effect (see chapter VI.2): The Hall conductivity stays constant as the Fermi energy increases
but then suddenly jumps by a discrete amount due to the contribution of the extended state
1This is an example of a quantum phase transition. The entire discussion is at zero temperature. In contrast
to the phase transition discussed in chapter V .3, here we vary Einstead of the temperature.
352 | VI. Field Theory and Condensed Matter
as the Fermi energy passes through E=0. Understanding this behavior quantitatively
poses a major challenge for condensed matter theorists. Indeed, many consider an analyticcalculation of the critical exponent νas one of the “Holy Grails” of condensed matter theory.
Green’s function formalism
So much for a lightning glimpse of localization theory. Fascinating though the localization
transition might be, what does quantum field theory have to do with it? This is after all afield theory text. Before proceeding we need a bit of formalism. Consider the so-calledGreen’s function G(z)≡/angbracketlefttr[1/(z−H)]/angbracketrightin the complex z-plane. Since tr[1 /(z−H)]=/summationtext
i1/(z−Ei), this function consists of a sum of poles at the eigenvalues Ei. Upon
averaging, the poles merge into a cut. Using the identity (I.2.13) lim
ε→0Im[1/(x+iε)]=
−πδ(x), we see that
ρ(E)=−1
πlim
ε→0ImG(E+iε) (1)
So if we know G(z) we know the density of states.
The infamous denominator
I can now explain how quantum field theory enters into the problem. We start by taking
the logarithm of the identity (A.15)
J†.K−1.J=log(/integraldisplay
Dϕ†Dϕe−ϕ†.K.ϕ+J†.ϕ+ϕ†.J)
(where as usual we have dropped an irrelevant term). Differentiating with respect to J†
andJand then setting J†andJequal to 0 we obtain an integral representation for the
inverse of a hermitean matrix:
(K−1)ij=/integraltext
Dϕ†Dϕe−ϕ†.K.ϕϕiϕ†
j/integraltext
Dϕ†Dϕe−ϕ†.K.ϕ(2)
(Incidentally, you may recognize this as essentially related to the formula (I.7.14) for the
propagator of a scalar field.) Now that we know how to represent 1 /(z−H)we have to take
its trace, which means setting i=jin (2) and summing. In our problem, H=− ∇2+V( x)
and the index icorresponds to the continuous variable xand the summation to an
integration over space. Replacing Kbyi(z−H) (and taking care of the appropriate delta
function) we obtain
tr−i
z−H=/integraldisplay
dDy⎧
⎨
⎩/integraltext
Dϕ†Dϕei/integraltext
dDx{∂ϕ†∂ϕ+[V( x)−z]ϕ†ϕ}ϕ(y)ϕ†(y)
/integraltext
Dϕ†Dϕei/integraltext
dDx{∂ϕ†∂ϕ+[V( x)−z]ϕ†ϕ}⎫
⎬
⎭(3)
VI.7. Disorder | 353
This is starting to look like a scalar field theory in D-dimensional Euclidean space with the
action S=/integraltext
dDx{∂ϕ†∂ϕ+[V( x)−z]ϕ†ϕ}. [Note that for (3) to be well defined zhas to be
in the lower half-plane.]
But now we have to average over V( x) , that is, integrate over Vwith the probability
distribution P(V) . We immediately run into the difficulty that confounded theorists for
a long time. The denominator in (3) stops us cold: If that denominator were not there,then the functional integration over V( x) would just be the Gaussian integral you have
learned to do over and over again. Can we somehow lift this infamous denominator intothe numerator, so to speak? Clever minds have come up with two tricks, known as thereplica method and the supersymmetric method, respectively. If you can come up withanother trick, fame and fortune might be yours.
Replicas
The replica trick is based on the well-known identity (1/x)=lim
n→0xn−1, which allows us to
write that much disliked denominator as
lim
n→0/parenleftBig/integraldisplay
Dϕ†Dϕei/integraltext
dDx{∂ϕ†∂ϕ+[V( x)−z]ϕ†ϕ}/parenrightBign−1
=lim
n→0/integraldisplayn/productdisplay
a=2Dϕ†
aDϕaei/integraltext
dDx/summationtextn
a=2{∂ϕ†
a∂ϕa+[V( x)−z]ϕ†
aϕa}
Thus (3) becomes
tr1
z−H=lim
n→0i/integraldisplay
dDy/integraldisplay/parenleftBiggn/productdisplay
a=1Dϕ†
aDϕa/parenrightBigg
ei/integraltext
dDx/summationtextn
a=1{∂ϕ†
a∂ϕa+[V( x)−z]ϕ†
aϕa}ϕ1(y)ϕ†
1(y) (4)
Note that the functional integral is now over ncomplex scalar fields ϕa. The field ϕhas
been replicated. For positive integers, the integrals in (4) are well defined. We hope thatthe limit n→0 will not blow up in our face.
Averaging over the random potential, we recover translation invariance; thus the in-
tegrand for/integraltext
d
Dydoes not depend on yand/integraltext
dDyjust produces the volume Vof the
system. Using (A.13) we obtain
/angbracketleftbigg
tr1
z−H/angbracketrightbigg
=iVlim
n→0/integraldisplay/parenleftBiggn/productdisplay
a=1Dϕ†
aDϕa/parenrightBigg
ei/integraltext
dDxLϕ1(0)ϕ†
1(0) (5)
where
L(ϕ)≡n/summationdisplay
a=1(∂ϕ†
a∂ϕa−zϕ†
aϕa)+ig2
2/parenleftBiggn/summationdisplay
a=1ϕ†
aϕa/parenrightBigg2
(6)
We obtain a field theory (with a peculiar factor of i)ofnscalar fields with a good old
ϕ4interaction invariant under O(n) (known as the replica symmetry.) Note the wisdom of
354 | VI. Field Theory and Condensed Matter
replacing Kbyi(z−H); if we didn’t include the ithe functional integral would diverge
at large ϕ, as you can easily check. For zin the upper half-plane we would replace Kby
−i(z−H). The quantity from which we can extract the desired averaged density of states
is given by the propagator of the scalar field. Incidentally, we can replace ϕ1(0)ϕ†
1(0)in (5)
by the more symmetric expression (1/n)/summationtextn
b=1ϕ†
bϕb.
Absorbing Vso that we are calculating the density of states per unit volume, we find
G(z)=ilim
n→0/integraldisplay/parenleftBiggn/productdisplay
a=1Dϕ†
aDϕa/parenrightBigg
eiS(ϕ)/parenleftBigg
1
nn/summationdisplay
b=1ϕ†
b(0)ϕb(0)/parenrightBigg
(7)
For positive integer nthe field theory is perfectly well defined, so the delicate step in the
replica approach is in taking the n→0 limit. There is a fascinating literature on this limit.
(Consult a book devoted to spin glasses.)
Some particle theorists used to speak disparagingly of condensed matter physics as dirt
physics, and indeed the influence of impurities and disorder on matter is one of the centralconcerns of modern condensed matter physics. But as we see from this example, in manyrespects there is no mathematical difference between averaging over randomness andsumming over quantum fluctuations. We end up with a ϕ
4field theory of the type that many
particle theorists have devoted considerable effort to studying in the past. Furthermore,Anderson’s surprising result that for D=2 any amount of disorder, no matter how small,
localizes all states means that we have to understand the field theory defined by (6) ina highly nontrivial way. The strength of the disorder shows up as the coupling g
2,s o
no amount of perturbation theory in g2can help us understand localization. Anderson
localization is an intrinsically nonperturbative effect.
Grassmannian approach
As I mentioned earlier, people have dreamed up not one, but two, tricks in dealing withthe nasty denominator. The second trick is based on what we learned in chapter II.5 onintegration over Grassmann variables: Let η(x) and¯η(x) be Grassmann fields, then
/integraldisplay
DηD¯ηe−/integraltext
d4x¯ηKη=CdetK=/tildewideC/parenleftbigg/integraldisplay
DϕDϕ†e−/integraltext
d4xϕ†Kϕ/parenrightbigg−1
where Cand/tildewideCare two uninteresting constants that we can absorb into the definition of
DηD¯η. With this identity we can write (3) as
tr1
z−H=i/integraldisplay
dDy/integraldisplay
Dϕ†DϕDηD ¯ηei/integraltext
dDx{{∂ϕ†∂ϕ+[V( x)−z]ϕ†ϕ}+{∂¯η∂η+[ V( x)−z]¯ηη}}ϕ(y)ϕ†(y) (8)
and then easily average over the disorder to obtain (per unit volume)
/angbracketleftbigg
tr1
z−H/angbracketrightbigg
=i/integraldisplay
Dϕ†DϕDηD ¯ηei/integraltext
dDxL(¯η,η,ϕ†,ϕ)ϕ(0)ϕ†(0) (9)
with
L(¯η,η,ϕ†,ϕ)=∂ϕ†∂ϕ+∂¯η∂η−z(ϕ†ϕ+¯ηη)+ig2
2(ϕ†ϕ+¯ηη)2(10)
VI.7. Disorder | 355
We end up with a field theory with bosonic (commuting) fields ϕ†andϕand fermionic
(anticommuting) fields ¯ηandηinteracting with a strength determined by the disorder. The
action Sexhibits an obvious symmetry rotating bosonic fields into fermionic fields and
vice versa, and hence this approach is known in the condensed matter physics communityas the supersymmetric method. (It is perhaps worth emphasizing that ¯ηandηare not
spinor fields, which we underline by not writing them as ¯ψandψ. The supersymmetry
here, perhaps better referred to as Grassmannian symmetry, is quite different from thesupersymmetry in particle physics to be discussed in chapter VIII.4.)
Both the replica and the supersymmetry approaches have their difficulties, and I was
not kidding when I said that if you manage to invent a new approach without some of thesedifficulties it will be met with considerable excitement by condensed matter physicists.
Probing localization
I have shown you how to calculate the averaged density of states ρ(E) . How do we study
localization? I will let you develop the answer in an exercise. From our earlier discussionit should be clear that we have to study an object obtained from (3) by replacing ϕ(y)ϕ
†(y)
byϕ(x)ϕ†(y)ϕ(y)ϕ†(x). If we choose to think of the replica field theory in the language of
particle physics as describing the interaction of some scalar meson, then rather pleasingly,we see that the density of states is determined by the meson propagator and localizationis determined by meson-meson scattering.
Exercises
VI.7.1 Work out the field theory that will allow you to study Anderson localization. [Hint: Consider the object
/angbracketleftbigg/parenleftbigg1
z−H/parenrightbigg
(x,y)/parenleftbigg1
w−H/parenrightbigg
(y,x)/angbracketrightbigg
for two complex numbers zandw. You will have to introduce two sets of replica fields, commonly denoted
byϕ+
aandϕ−
a.] {Notation: [1 /(z−H)](x,y)denotes the xyelement of the matrix or operator [1 /(z−H)].}
VI.7.2 As another example from the literature on disorder, consider the following problem. Place Npoints
randomly in a D-dimensional Euclidean space of volume V. Denote the locations of the points by
/vectorxi(i=1 ,..., N). Let
f(/vectorx)=(−)/integraldisplaydDk
(2π)Dei/vectork/vectorx
k2+m2
Consider the NbyNmatrix Hij=f/parenleftbig
/vectorxi−/vectorxj/parenrightbig
. Calculate ρ(E) , the density of eigenvalues of Has we
average over the ensemble of matrices, in the limit N→∞ ,V→∞ , with the density of points ρ≡N/V
(not to be confused with ρ(E) of course) held fixed. [Hint: Use the replica method and arrive at the field
theory action
S(ϕ)=/integraldisplay
dDx/bracketleftBiggn/summationdisplay
a=1(|∇ϕa|2+m2|ϕa|2)−ρe−(1/z)/summationtextn
a=1|ϕa|2/bracketrightBigg
This problem is not entirely trivial; if you need help consult M. M ´ezard et al., Nucl. Phy. B559: 689, 2000,
cond-mat/9906135.
VI.8Renormalization Group Flow as a Natural Concept
in High Energy and Condensed Matter Physics
Therefore, conclusions based on the renormalization
group arguments...a r e dangerous and must be viewed
with due caution. So is it with all conclusions from localrelativistic field theories.
—J. Bjorken and S. Drell, 1965
It is not dangerous
The renormalization group represents the most important conceptual advance in quantumfield theory over the last three or four decades. The basic ideas were developed simultane-ously in both the high energy and condensed matter physics communities, and in someareas of research renormalization group flow has become part of the working language.
As you can easily imagine, this is an immensely rich and multifaceted subject, which we
can discuss from many different points of view, and a full exposition would require a bookin itself. Unfortunately, there has never been a completely satisfactory and comprehensivetreatment of the subject. The discussions in some of the older books are downrightmisleading and confused, such as the well-known text from which I learned quantum fieldtheory and from which the quote above was taken. In the limited space available here, Iwill attempt to give you a flavor of the subject rather than all the possible technical details.I will first approach it from the point of view of high energy physics and then from thatof condensed matter physics. As ever, the emphasis will be on the conceptual rather thanthe computational. As you will see, in spite of the order of my presentation, it is easierto grasp the role of the renormalization group in condensed matter physics than in highenergy physics.
I laid the foundation for the renormalization group in chapter III.1—I do plan ahead!
Let us go back to our experimentalist friend with whom we were discussing λϕ
4theory.
We will continue to pretend that our world is described by a simple λϕ4theory and that an
approximation to order λ2suffices.
VI.8. Renormalization Group Flow | 357
What experimentalists insist on
Our experimentalist friend was not interested in the coupling constant λwe wrote down
on a piece of paper, a mere Greek letter to her. She insisted that she would accept onlyquantities she and her experimental colleagues can actually measure, even if only inprinciple. As a result of our discussion with her we sharpened our understanding of whata coupling constant is and learned that we should define a physical coupling constant by[see (III.1.4)]
λP(μ)=λ−3Cλ2log/parenleftbigg/Lambda12
μ2/parenrightbigg
+O(λ3) (1)
At her insistence, we learned to express our result for physical amplitudes in terms of
λP(μ), and not in terms of the theoretical construct λ. In particular, we should write the
meson-meson scattering amplitude as
M=−iλP(μ)+iCλP(μ)2/bracketleftbigg
log/parenleftbiggμ2
s/parenrightbigg
+log/parenleftbiggμ2
t/parenrightbigg
+log/parenleftbiggμ2
u/parenrightbigg/bracketrightbigg
+O[λP(μ)3] (2)
What is the physical significance of λP(μ)? T o be sure, it measures the strength of
the interaction between mesons as reflected in (2). But why one particular choice of μ?
Clearly, from (2) we see that λP(μ) is particularly convenient for studying physics in the
regime in which the kinematic variables s,t, and uare all of order μ2. The scattering
amplitude is given by −iλP(μ) plus small logarithmic corrections. (Recall from a footnote
in chapter III.3 that the renormalization point s0=t0=u0=μ2is adopted purely for
theoretical convenience and cannot be reached in actual experiments. For our conceptualunderstanding here this is not a relevant issue.) In short, λ
P(μ) is known as the coupling
constant appropriate for physics at the energy scale μ.
In contrast, if we were so idiotic as to use the coupling constant λP(μ/prime)while exploring
physics in the regime with s,t, anduof order μ2, with μ/primevastly different from μ, then we
would have a scattering amplitude
M=−iλP(μ/prime)+iCλP(μ/prime)2/bracketleftbigg
log/parenleftbiggμ/prime2
s/parenrightbigg
+log/parenleftbiggμ/prime2
t/parenrightbigg
+log/parenleftbiggμ/prime2
u/parenrightbigg/bracketrightbigg
+O[λP(μ/prime)3] (3)
in which the second term [with log (μ/prime2/μ2)large] can be comparable to or larger than the
first term. The coupling constant λP(μ/prime)is not a convenient choice. Thus, for each energy
scaleμthere is an “appropriate” coupling constant λP(μ).
Subtracting (2) from (3) we can easily relate λP(μ) andλP(μ/prime)forμ∼μ/prime:
λP(μ/prime)=λP(μ)+3CλP(μ)2log/parenleftbiggμ/prime2
μ2/parenrightbigg
+O[λP(μ)3] (4)
We can express this as a differential “flow equation”
μd
dμλP(μ)=6CλP(μ)2+O(λ3
P) (5)
358 | VI. Field Theory and Condensed Matter
As you have already seen repeatedly, quantum field theory is full of historical misnomers.
The description of how λP(μ) changes with μis known as the renormalization group.
The only appearance of a group concept here is the additive group of transformationμ→μ+δμ.
For the conceptual discussion in chapter III.1 and here, we don’t need to know what
the constant Chappens to be. If Chappens to be negative, then the coupling λ
P(μ) will
decrease as the energy scale μincreases, and the opposite will occur if Chappens to be
positive. (In fact, the sign is positive, so that as we increase the energy scale, λPflows away
from the origin.)
Flow of the electromagnetic coupling
The behavior of λis typical of coupling constants in 4-dimensional quantum field theories.
For example, in quantum electrodynamics, the coupling eor equivalently α=e2/4π,
measures the strength of the electromagnetic interaction. The story is exactly as that toldfor the λϕ
4theory: Our experimentalist friend is not interested in the Latin letter e, but
wants to know the actual interaction when the relevant momenta squared are of the orderμ
2. Happily, we have already done the computation: We can read off the effective coupling
at momentum transferred squared q2=μ2from (III.7.14):
eP(μ)2=e2 1
1+e2/Pi1(μ2)/similarequale2[1−e2/Pi1(μ2)+O(e4)]
T akeμmuch larger than the electron mass mbut much smaller than the cutoff mass M.
Then from (III.7.13)
μd
dμeP(μ)=−1
2e3μd
dμ/Pi1(μ2)+O(e5)=+1
12π2e3
P+O(e5
P) (6)
We learn that the electromagnetic coupling increases as the energy scale increases.
Electromagnetism becomes stronger as we go to higher energies, or equivalently shorterdistances.
Physically, the origin of this phenomenon is closely related to the physics of dielectrics.
Consider a photon interacting with an electron, which we will call the test electron toavoid confusion in what follows. Due to quantum fluctuations, as described way back inchapter I.1, spacetime is full of electron-positron pairs, popping in and out of existence.Near the test electron, the electrons in these virtual pairs are repelled by the test electronand thus tend to move away from the test electron while the positrons tend to move towardthe test electron. Thus, at long distances, the charge of the test electron is shielded to someextent by the cloud of positrons, causing a weaker coupling to the photon, while at shortdistances the coupling to the photon becomes stronger. The quantum vacuum is just asmuch a dielectric as a lump of actual material.
You may have noticed by now that the very name “coupling constant” is a terrible
misnomer due to the fact that historically much of physics was done at essentially oneenergy scale, namely “almost zero”! In particular, people speak of the fine structure
VI.8. Renormalization Group Flow | 359
“constant” α=1/137 and crackpots continue to try to “derive” the number 137 from
numerology or some fancier method. In fact, αis merely the coupling “constant” of the
electromagnetic interaction at very low energies. It is an experimental fact that α, more
properly written as αP(μ)≡e2
P(μ)/4π, varies with the energy scale μwe are exploring.
But alas, we are probably stuck with the name “coupling constant.”
Renormalization group flow
In general, in a quantum field theory with a coupling constant g, we have the renormal-
ization group flow equation
μdg
dμ=β(g) (7)
which is sometimes written as dg/dt =β(g) upon defining t≡log(μ/μ 0). I will now
suppress the subscript Pon physical coupling constants. If the theory happens to have
several coupling constants gi,i=1,... ,N, then we have
dgi
dt=βi(g1,... ,gN) (8)
We can think of (g1,... ,gN)as the coordinate of a particle in N-dimensional space,
tas time, and βi(g1,... ,gN)a position dependent velocity field. As we increase μort
we would like to study how the particle moves or flows. For notational simplicity, we willnow denote (g
1,... ,gN)collectively as g. Clearly, those couplings at which βi(g∗)(for all
i)happen to vanish are of particular interest: g∗is known as a fixed point. If the velocity
field around a fixed point g∗is such that the particle moves toward that point (and once
reaching it stays there since its velocity is now zero) the fixed point is known as attractive orstable. Thus, to study the asymptotic behavior of a quantum field theory at high energieswe “merely” have to find all its attractive fixed points under the renormalization groupflow. In a given theory, we can typically see that some couplings are flowing toward largervalues while others are flowing toward zero.
Unfortunately, this wonderful theoretical picture is difficult to implement in practice
because we essentially have no way of calculating the functions β
i(g). In particular, g∗could
well be quite large, associated with what is known as a strong coupling fixed point, and
perturbation theory and Feynman diagrams are of no use in determining the propertiesof the theory there. Indeed, we know the fixed point structure of very few theories.
Happily, we know of one particularly simple fixed point, namely g
∗=0, at which
perturbation theory is certainly applicable. We can always evaluate (8) perturbatively:
dgi/dt=cjk
igjgk+djkl
igjgkgl+... . (In some theories, the series starts with quadratic
terms and in others, with cubic terms. Sometimes there is also a linear term.) Thus, as we
have already seen in a couple of examples, the asymptotic or high energy behavior of the
theory depends on the sign of βiin (8).
Let us now join the film “Physics History” already in progress. In the late 1960s,
experimentalists studying the so-called deep inelastic scattering of electrons on protons
360 | VI. Field Theory and Condensed Matter
discovered that their data seemed to indicate that after being hit by a highly energetic
electron, one of the quarks inside the proton would propagate freely without interactingstrongly with the other quarks. Normally, of course, the three quarks inside the proton arestrongly bound to each other to form the proton. Eventually, a few theorists realized thatthis puzzling state of affairs could be explained if the theory of strong interaction is suchthat the coupling flows toward the fixed point g
∗=0. If so, then the strong interaction
between quarks would actually weaken at higher and higher energy scales.
All of this is of course now “obvious” with the benefit of hindsight, but dear students,
remember that at that time field theory was pronounced as possibly unsuitable for youngminds and the renormalizable group was considered “dangerous” even in a field theorytext!
The theory of strong interaction was unknown. But if we were so bold as to accept the
dangerous renormalization group ideas then we might even find the theory of the stronginteraction by searching for asymptotically free theories, which is what theories with anattractive fixed point at g
∗=0 became known as.
Asymptotically free theories are clearly wonderful. Their behavior at high energies can
be studied using perturbative methods. And so in this way the fundamental theory of thestrong interaction, now known as quantum chromodynamics, about which more later, wasfound.
Looking at physics on different length scales
The need for renormalization groups is really transparent in condensed matter physics.Instead of generalities, let me focus on a particularly clear example, namely surfacegrowth. Indeed, that was why I chose to introduce the Kardar-Parisi-Zhang equation inchapter VI.6. We learned that to study surface growth we have to evaluate the functionalor path integral
Z(/Lambda1) =/integraldisplay
/Lambda1Dhe−S(h). (9)
with, you will recall,
S(h)=1
2/integraldisplay
dD/vectorxd t/parenleftbigg∂h
∂t−∇2h−g
2(∇h)2/parenrightbigg2
. (10)
This defines a field theory. As with any field theory, and as I indicate, a cutoff /Lambda1has to be
introduced. We integrate over only those field configurations h(/vectorx,t)that do not contain
Fourier components with /vectorkandωlarger than /Lambda1. (In principle, since this is a nonrelativistic
theory we should have different cutoffs for /vectorkand for ω, but for simplicity of exposition let
us just refer to them together generically as /Lambda1.)The appearance of the cutoff is completely
physical and necessary. At the very least, on length scales comparable to the size of therelevant molecules, the continuum description in terms of the field h(/vectorx,t)has long since
broken down.
Physically, since the random driving term η(/vectorx,t)is a white noise, that is, ηat/vectorxand
at/vectorx
/prime(and also at different times) are not correlated at all, we expect the surface to look
VI.8. Renormalization Group Flow | 361
(a)
(b)
Figure VI.8.1
very uneven on a microscopic scale, as depicted in figure VI.8.1a. But suppose we are not
interested in the detailed microscopic structure, but more in how the surface behaves ona larger scale. In other words, we are content to put on blurry glasses so that the surfaceappears as in figure VI.8.1b. This is a completely natural way to study a physical system,one that we are totally familiar with from day one in studying physics. We may be interestedin physics over some length scale Land do not care about what happens on length scales
much less than L.
The renormalization group is the formalism that allows us to relate the physics on differ-
ent length scales or, equivalently, physics on different energy scales. In condensed matterphysics, one tends to think of length scales, and in particle physics, of energy scales. Themodern approach to renormalization groups came out of the study of critical phenomenaby Kadanoff, Fisher, Wilson, and others, as mentioned in chapter V .3. Consider, for exam-ple, the Ising model, with the spin at each site either up or down and with a ferromagneticinteraction between neighboring spins. At high temperatures, the spins point randomlyup and down. As the temperature drops toward the ferromagnetic transition point, islandsof up spins (we say up spins to be definite, we could just as easily talk of down spins) startto appear. They grow ever larger in size until the critical temperature T
cat which all the
spins in the entire system point up. The characteristic length scale of the physics at any
particular temperature is given by the typical size of the islands. The physically motivatedblock spin method of Kadanoff et al. treats blocks of up spin as one single effective upspin, and similarly blocks of down spins. The notion of a renormalization group is thenthe natural one for describing these effective spins by an effective Hamiltonian appropriateto that length scale.
It is more or less clear how to implement this physical idea of changing length scales in
the functional integral (9). We are supposed to integrate over those h(/vectork,ω), with /vectorkandω
less than /Lambda1. Suppose we do only a fraction of what we are supposed to do. Let us integrate
362 | VI. Field Theory and Condensed Matter
over those h(/vectork,ω)with/vectorkandωlarger than /Lambda1−δ/Lambda1 but smaller than /Lambda1. This is precisely
what we mean when we say that we don’t care about the fluctuations of h(/vectorx,t)on length
and time scales less than (/Lambda1 −δ/Lambda1)−1.
Putting on blurry glasses
For the sake of simplicity, let us go back to our favorite, the λϕ4theory, instead of the surface
growth problem. Recall from the preceding chapters the importance of the Euclidean λϕ4
theory in modern condensed matter theory. So, continue the λϕ4theory to Euclidean space
and stare at the integral
Z(/Lambda1) =/integraldisplay
/Lambda1Dϕe−/integraltext
ddxL(ϕ)(11)
The notation/integraltext
/Lambda1instructs us to include only those field configurations ϕ(x)=/integraltext
[ddk/(2π)d]eikxϕ(k) such that ϕ(k)=0 for|k|≡(/summationtextd
i=1k2
i)1
2larger than /Lambda1. As explained
in the text this amounts to putting on blurry glasses with resolution L=1//Lambda1:W ed on o t
admit or see fluctuations with length scales less than L.
Evidently, the O(d) invariance, namely the Euclidean equivalent of Lorentz invariance,
will make our lives considerably easier. In contrast, for the surface growth problem we willneed special glasses that blur space and time differently.
1
We are now ready to make our glasses blurrier by letting /Lambda1→/Lambda1−δ/Lambda1 (withδ/Lambda1 > 0).
Write ϕ=ϕs+ϕw(sfor “smooth” and wfor “wriggly”), defined such that the Fourier
components ϕs(k) andϕw(k) are nonzero only for |k|≤(/Lambda1 −δ/Lambda1) and (/Lambda1 −δ/Lambda1)≤|k|≤
/Lambda1, respectively. (Obviously, the designation “smooth” and “wriggly” is for convenience.)Plugging into (11) we can write
Z(/Lambda1) =/integraldisplay
/Lambda1−δ/Lambda1Dϕse−/integraltext
ddxL(ϕs)/integraldisplay
Dϕwe−/integraltext
ddxL1(ϕs,ϕw)(12)
where all the terms in L1(ϕs,ϕw)depend on ϕw. (What we are doing here is somewhat
reminiscent of what we did in chapter IV .3.) Imagine doing the integral over ϕw. Call the
result
e−/integraltext
ddxδL(ϕs)≡/integraldisplay
Dϕwe−/integraltext
ddxL1(ϕs,ϕw)
and thus we have
Z(/Lambda1) =/integraldisplay
/Lambda1−δ/Lambda1Dϕse−/integraltext
ddx[L(ϕs)+δL(ϕs)](13)
There, we have done it! We have rewritten the theory in terms of the “smooth” field ϕs.
Of course, this is all formal, since in practice the integral over ϕwcan only be done
perturbatively assuming that the relevant couplings are small. If we could do the integral
1In condensed matter physics, the so-called dynamical exponent zmeasures this difference. More precisely,
in the context of the surface growth problem, the correlator (introduced in chapter VI.6) satisfies the dynamicscaling form given in (VI.6.3). Naively, the dynamical exponent zshould be 2. (For a brief review of all this, see
M. Kardar and A. Zee, Nucl. Phys. B464[FS]: 449, 1996, cond-mat/9507112.)
VI.8. Renormalization Group Flow | 363
overϕwexactly, we might as well just do the integral over ϕand then we would have no
need for all this renormalization group stuff.
For pedagogical purposes, consider more generally L=1
2(∂ϕ)2+/summationtext
nλnϕn+... (so
thatλ2is the usual1
2m2andλ4the usual λ.)Since terms such as ∂ϕs∂ϕwintegrate to zero,
we have
/integraldisplay
ddxL1(ϕs,ϕw)=/integraldisplay
ddx/parenleftbigg1
2(∂ϕw)2+1
2m2ϕ2
w+.../parenrightbigg
withϕshiding in the (...). This describes a field ϕwinteracting with both itself and a
background field ϕs(x). By symmetry considerations δL(ϕs)has the same form as L(ϕs)
but with different coefficients. Adding δL(ϕs)toL(ϕs)thus shifts2the couplings λn[and the
coefficient of1
2(∂ϕs)2.] These shifts generate the flow in the space of couplings I described
earlier.
We could have perfectly well left (13) as our end result. But suppose we want to compare
(13) with (11). Then we would like to change the/integraltext
/Lambda1−δ/Lambda1in (13) to/integraltext
/Lambda1. For convenience,
introduce the real number b< 1b y/Lambda1−δ/Lambda1=b/Lambda1.I n/integraltext
/Lambda1−δ/Lambda1we are told to integrate over
fields with |k|≤b/Lambda1 . So all we have to do is make a trivial change of variable: Let k=bk/prime
so that |k/prime|≤/Lambda1 . But then correspondingly we have to change x=x/prime/bso that eikx=eik/primex/prime.
Plugging in, we obtain
/integraldisplay
ddxL(ϕs)=/integraldisplay
ddx/primeb−d/bracketleftBigg
1
2b2(∂/primeϕs)2+/summationdisplay
nλnϕn
s+.../bracketrightBigg
(14)
where ∂/prime=∂/∂x/prime=(1/b)∂/∂x . Define ϕ/primebyb2−d(∂/primeϕs)2=(∂/primeϕ/prime)2or in other words ϕ/prime=
b1
2(2−d)ϕs. Then (14) becomes
/integraldisplay
ddx/prime/bracketleftBigg
1
2(∂/primeϕ/prime)2+/summationdisplay
nλnb−d+(n/2 )(d−2)ϕ/primen+.../bracketrightBigg
Thus, if we define the coefficient of ϕ/primenasλ/prime
nwe have
λ/prime
n=b(n/2)(d−2)−dλn (15)
an important result in renormalization group theory.
Relevant, irrelevant, and marginal
Let us absorb what this means. (For the time being, let us ignore δL(ϕs)to keep the
discussion simple.) As we put on blurrier glasses, in other words, as we become interestedin physics over longer distance scales, we can once again write Z(/Lambda1) as in (11) except that
the couplings λ
nhave to be replaced by λ/prime
n. Since b< 1 we see from (15) that the λn’s with
(n/2)(d−2)−d> 0 get smaller and smaller and can eventually be neglected. A dose of
jargon here: The corresponding operators ϕn(for historical reasons we revert for an instant
2T erms such as (∂ϕ)4can also be generated and that is why I wrote L(ϕ) with the (...)under which terms such
as these can be swept. You can check later that for most applications these terms are irrelevant in the technicalsense to be defined below.
364 | VI. Field Theory and Condensed Matter
from the functional integral language to the operator language) are called irrelevant. They
are the losers. Conversely, the winners, namely the ϕn’s for which (n/2)(d−2)−d< 0,
are called relevant. Operators for which (n/2)(d−2)−d=0 are called marginal.
For example, take n=2:m/prime2=b−2m2and the mass term is always relevant in any
dimension. On the other hand, take n=4, and we see that λ/prime=bd−4λandϕ4is relevant
ford< 4, irrelevant for d> 4, and marginal at d=4. Similarly, λ/prime
6=b2d−6λandϕ6is
marginal at d=3 and becomes irrelevant for d> 3.
We also see that d=2 is special: All the ϕn’s are relevant.
Now all this may ring a bell if you did the exercises religiously. In exercise III.2.1
you showed that the coupling λnhas mass dimension [ λn]=(n/2)(2−d)+d. Thus, the
quantity (n/2)(d−2)−dis just the length dimension of λn. For example, for d=4,
λ6has mass dimension −2 and thus as explained in chapter III.2 the ϕ6interaction is
nonrenormalizable, namely that it has nasty behavior at high energy. But condensed matterphysicists are interested in the long distance limit, the opposite limit from the one thatinterests particle physicsts. Thus, it is the nasty guys like ϕ
6that become irrelevant in the
long distance limit.
One more piece of jargon: Given a scalar field theory, the dimension dat which the most
relevant interaction becomes marginal is known as the critical dimension in condensedmatter physics. For example, the critical dimension for a ϕ
6theory is 3. It is now just
a matter of “high school arithmetic” to translate (15) into differential form. Write λ/prime
n=
λn+δλn; then from b=1−(δ/Lambda1//Lambda1) we have δλn=− [n
2(d−2)−d]λn(δ/Lambda1//Lambda1) .
Let us now be extra careful about signs. As I have already remarked, for (n/2)(d−2)−
d> 0 the coupling λn(which, for definiteness, we will think of as positive) get smaller,
as is evident from (15). But since we are decreasing /Lambda1to/Lambda1−δ/Lambda1, a positive δ/Lambda1 actually
corresponds to the resolution of our blurry glasses L=/Lambda1−1changing to L+L(δ/Lambda1//Lambda1) .
Thus we obtain
Ldλn
dL=−/bracketleftbiggn
2(d−2)−d/bracketrightbigg
λn, (16)
so that for (n/2)(d−2)−d> 0 a positive λnwould decrease as Lincreases.3
In particular, for n=4,L(dλ/dL) =(4−d)λ . In most condensed matter physics appli-
cations, d≤3 and so λincreases as the length scale of the physics under study increases.
Theϕ4coupling is relevant as noted above.
TheδL(ϕs), which we provisionally neglected, contributes an additional term, which we
call dynamical in contrast to the geometrical or “trivial” term displayed, to the right-hand
side of (16). Thus, in general L(dλn/dL)=− [(n/2)(d−2)−d]λn+K(d ,n,...,λj,...),
with the dynamical term Kdepending not only on dandn, but also on all the other
couplings. [For example, in (5) the “trivial” term vanishes since we are in 4-dimensionalspacetime; there is only a dynamical contribution.]
3Note that what appears on the right-hand side is minus the length dimension of λn, not the length dimension
(n/2)(d−2)−das one might have guessed naively.
VI.8. Renormalization Group Flow | 365
As you can see from this discussion, a more descriptive name for the renormalization
group might be “the trick of doing an integral a little bit at a time.”
Exploiting symmetry
T o determine the renormalization group flow of the coupling gin the surface growth
problem we can repeat the same type of computation we did to determine the flow of thecoupling λand of ein our two previous examples, namely we would calculate, to use the
language of particle physics, the amplitude for h-h scattering to one loop order. But instead,
let us follow the physical picture of Kadanoff et al. In Z(/Lambda1)=/integraltext
Dhe
−S(h)we integrate over
only those h(/vectork,ω)with/vectorkandωlarger than /Lambda1−δ/Lambda1 but less than /Lambda1.
I will now show you how to exploit the symmetry of the problem to minimize our labor.
The important thing is not necessarily to learn about the dynamics of surface growth,but to learn the methodology that will serve you well in other situations. I have pickeda particularly “difficult” nonrelativistic problem whose symmetries are not manifest, sothat if you master the renormalization group for this problem you will be ready for almostanything.
Imagine having done this partial integration and call the result/integraltext
Dhe
−˜S(h). At this point
you should work out the symmetries of the problem as indicated in the exercises. Thenyou can argue that ˜S(h) must have the form
˜S(h)=1
2/integraldisplay
dDxd t/bracketleftbigg/parenleftbigg
α∂
∂t−β∇2/parenrightbigg
h−αg
2(∇h)2/bracketrightbigg2
+... , (17)
depending on two parameters αandβ. The ( ...)indicates terms involving higher powers
ofhand its derivatives. The simplifying observation is that the same coefficient αmultiplies
both∂h/∂t and(g/2)(∇h)2. Once we know αandβthen by suitable rescaling we can
bring the action ˜S(h) back into the same form as S(h) and thus find out how gchanges.
Therefore, it suffices to look at the (∂h/∂t)2and(∇2h)2terms in the action, or equivalently
at the propagator, which is considerably simpler to calculate. As Rudolf Peierls once said4
to the young Hans Bethe, “Erst kommt das Denken, dann das Integral.” (Roughly, “Firstthink, then do the integral.”) We will not do the computation here. Suffice it to note that g
has the high school dimension of (length)
1
2(D−2)(see exercise VI.8.5). Thus, according to
the preceding discussion we should have
Ldg
dL=1
2(2−D)g+cDg3+... (18)
A detailed calculation is needed to determine the coefficient cD, which obviously
depends on the dimension of space Dsince the Feynman integrals depend on D.
The equation tells us how g, an effective measure of nonlinearity in the physics of
4John Wheeler gave me similar advice when I was a student: “Never calculate without first knowing the
answer.”
366 | VI. Field Theory and Condensed Matter
surface growth, changes when we change the length scale L. For the record, cD=
[S(D)/ 4(2π)D](2D−3)/D , with S(D) theD-dimensional solid angle. The interesting
factor is of course (2D−3), changing sign between5D=1 and 2.
Localization
As I said earlier, renormalization group flow has literally become part of the language of
condensed matter and high energy physics. Let me give you another example of the powerof the renormalization group. Go back to Anderson localization (chapter VI.7), which soastonished the community at the time. People were surprised that the localization behaviordepends so drastically on the dimension of space D, and perhaps even more so, that for
D=2 all states are localized no matter how weak the strength of the disorder. Our usual
physical intuition would say that there is a critical strength. As we will now see, bothfeatures are quite naturally accounted for in the renormalization group language. Already,you see in (18) that Denters in an essential way.
I now offer you a heuristic but beautiful (at least to me) argument given by Abra-
hams, Anderson, Licciardello, and Ramakrishnan, who as a result became known to thecondensed matter community as the “Gang of Four.” First, you have to understand thedifference between conductivity σand conductance Gin solid state physics lingo. Conduc-
tivity
6is defined by /vectorJ=σ/vectorE, where /vectorJmeasures the number of electrons passing through
a unit area per unit time. Conductance Gis the inverse of resistance (the mnemonic: the
two words rhyme). Resistance Ris the property of a lump of material and defined in high
school physics by V=IR, where the current Imeasures the number of electrons passing
by per unit time. T o relate σandG, consider a lump of material, taken to be a cube of size
L, with a voltage drop Vacross it. Then I=JL2=σEL2=σ(V/L)L2=σLV and thus7
G(L)=1/R=I/V=σL. Next, let us go to two dimensions. Consider a thin sheet of ma-
terial of length and width Land thickness a/lessmuchL. (We are doing real high school physics,
not talking about some sophisticated field theorist’s idea of two dimensional space!) Again,apply a voltage drop Vover the length L:I=J(aL) =σEaL =σ(V/L)aL =σVa and so
G(L)=1/R=I/V=σa. I will let you go on to one dimension: Consider a wire of length
Land width and thickness a. In this way, we obtain G(L)∝L
D−2. Incidentally, condensed
matter physicists customarily define a dimensionless conductance g(L)≡/planckover2piG(L)/e2.
5Incidentally, the theory is exactly solvable for D=1 (with methods not discussed in this book).
6Over the years I have asked a number of high energy theorists how is it possible to obtain /vectorJ=σ/vectorE, which
manifestly violates time reversal invariance, if the microscopic physics of an electron scattering on an impurityatom perfectly well respects time reversal invariance. Very few knew the answer. The resolution of this apparentparadox is in the order of limits! Condensed matter theorists calculate a frequency and wave vector dependent
conductivity σ(ω ,/vectork)and then take the limit ω,/vectork→0 and /vectork
2/ω→0. Before the limit is taken, time reversal
invariance holds. The time it takes the particle to find out that it is in a box of size of order 1 /kis of order 1 /(Dk2)
(withDthe diffusion constant). The physics is that this time has to be much longer than the observation time
∼1/ω.
7Sam T reiman told me that when he joined the U.S. Army as a radio operator he was taught that there were
three forms of Ohm’s law: V=IR,I=V/R , andR=V/I . In the second equality here we use the fourth form.
VI.8. Renormalization Group Flow | 367
1
−1β(g)
gcgD = 3
D = 2
D = 1
Figure VI.8.2
We also know the behavior of g(L) when g(L) is small or, in other words, when the
material is an insulator for which we expect g(L)∼ce−L/ξ, with ξsome length charac-
teristic of the material and determined by the microscopic physics. Thus, for g(L) small,
L(dg/dL) =−(L/ξ)g(L) =g(L)[log g(L)−logc], where the constant log cis negligible
in the regime under consideration.
Putting things together, we obtain
β(g)≡L
gdg
dL=/braceleftBigg(D−2)+... for large g
logg+... for small g(19)
First, a trivial note: in different subjects, people define β(g) differently (without affecting
the physics of course). In localization theory, β(g) is traditionally defined as dlogg/d logL
as indicated here. Given (19) we can now make a “most plausible” plot of β(g) as shown
in figure VI.8.2. You see that for D=2 (and D=1)the conductance g(L) always flows
toward 0 as we go to long distances (macroscopic measurements on macroscopic materials)regardless of where we start. In contrast, for D=3, ifg
0the initial value of gis greater than
a critical gctheng(L) flows to infinity (presumably cut off by physics we haven’t included)
and the material is a metal, while if g0<gc, the material is an insulator. Incidentally,
condensed matter theorists often speak of a critical dimension Dcat which the long
distance behavior of a system changes drastically; in this case Dc=2.
Effective description
In a sense, the renormalization group goes back to a basic notion of physics, that the
effective description can and should change as we move from one length scale to another.For example, in hydrodynamics we do not have to keep track of the detailed interaction
368 | VI. Field Theory and Condensed Matter
among water molecules. Similarly, when we apply the renormalization group flow to the
strong interaction, starting at high energies and moving toward low energies, the effectivedescription goes from a theory of quarks and gluons to a theory of nucleons and mesons.In this more general picture then, we no longer think of flowing in a space of couplingconstants, but in “the space of Hamiltonians” that some condensed matter physicists liketo talk about.
Exercises
VI.8.1 Show that the solution of dg/dt =−bg3+...is given by
1
α(t)=1
α(0)+8πbt+... (20)
where we defined α(t)=g(t)2/4π.
VI.8.2 In our discussion of the renormalization group, in λϕ4theory or in QED, for the sake of simplicity
we assumed that the mass mof the particle is much smaller than μand thus set mequal to zero. But
nothing in the renormalization group idea tells us that we can’t flow to a mass scale below m. Indeed, in
particle physics many orders of magnitude separate the top quark mass mtfrom the up quark mass mu.
We might want to study how the strong interaction coupling flows from some mass scale far above mt
down to some mass scale μbelow mtbut still large compared to mu. As a crude approximation, people
often set any mass mbelow μequal to zero and any mabove μto infinity (i.e., not contributing to the
renormalization group flow). In reality, as μapproaches mfrom above the particle starts to contribute
less and drops out as μbecomes much less than m. T aking either the λϕ4theory or QED study this
so-called threshold effect.
VI.8.3 Show that (10) is invariant under the so-called Galilean transformation
h(/vectorx,t)→h/prime(/vectorx,t)=h/parenleftbig
/vectorx+g/vectorut,t/parenrightbig
+/vectoru./vectorx+g
2u2t (21)
Show that because of this symmetry only two parameters αandβappear in (17).
VI.8.4 In˜S(h) only derivatives of the field hcan appear and not the field itself. (Since the transformation
h(/vectorx,t)→h(/vectorx,t)+cwithca constant corresponds to a trivial shift of where we measure the surface
height from, the physics must be invariant under this transformation.) T erms involving only one power
ofhcannot appear since they are all total divergences. Thus, ˜S(h) must start with terms quadratic in h.
Verify that the ˜S(h) given in (17) is indeed the most general. A term proportional to (∇h)2is also allowed
by symmetries and is in fact generated. However, such a term can be eliminated by transforming to amoving coordinate frame h→h+ct.
VI.8.5 Show that ghas the high school dimension of (length)1
2(D−2). [Hint: The form of S(h) implies that thas
the dimension of length squared and so hhas the dimension (length)1
2(2−D).] Comparing the terms ∇2h
andg(∇h)2we determine the dimension of g.]
VI.8.6 Calculate the hpropagator to one loop order. Extract the coefficients of the ω2andk4terms in a low
frequency and wave number expansion of the inverse propagator and determine αandβ.
VI.8.7 Study the renormalization group flow of gforD=1, 2, 3.
Part VII Grand Unification
This page intentionally left blank
VII.1Quantizing Yang-Mills Theory
and Lattice Gauge Theory
One reason that Y ang-Mills theory was not immediately taken up by physicists is that
people did not know how to calculate with it. At the very least, we should be able towrite down the Feynman rules and calculate perturbatively. Feynman himself took up thechallenge and concluded, after looking at various diagrams, that extra fields with ghostlikeproperties had to be introduced for the theory to be consistent. Nowadays we know howto derive this result more systematically.
The story goes that Feynman wanted to quantize gravity but Gell-Mann suggested to
him to first quantize Y ang-Mills theory as a warm-up exercise.
Consider pure Y ang-Mills theory—it will be easy to add matter fields later. Follow what
we have learned. Split the Lagrangian L=L
0+L1as usual into two pieces (we also choose
to scale A→gA) :
L0=−1
4(∂μAa
ν−∂νAa
μ)2(1)
and
L1=−1
2g(∂μAaν−∂νAa
μ)fabcAbμAcν−1
4g2fabcfadeAbμAcνAdμAeν(2)
Then invert the differential operator in the quadratic piece (1) to obtain the propagator.
This part looks the same as the corresponding procedure for quantum electrodynamics,except for the occurrence of the index a. Just as in electrodynamics, the inverse does not
exist and we have to fix a gauge.
I built up the elaborate Faddeev-Popov method to quantize quantum electrodynamics
and as I noted, it was a bit of overkill in that context. But here comes the payoff: We can nowturn the crank. Recall from chapter III.4 that the Faddeev-Popov method would give us
Z=/integraldisplay
DAeiS(A)/Delta1(A)δ [f (A)] (3)
with/Delta1(A) ≡{/integraltext
Dgδ [f( Ag)]}−1andS(A)=/integraltext
d4xLthe Y ang-Mills action. (As in chap-
ter III.4, Ag≡gAg−1−i(∂g)g−1denotes the gauge transform of A. Here g≡g(x) denotes
372 | VII. Grand Unification
the group element that defines the gauge transformation at xand is obviously not to be
confused with the coupling constant.)
Since /Delta1(A) appears in (3) multiplied by δ[f (A)], in the integral over gwe expect, for
a reasonable choice of f (A) , only infinitesimal gto be relevant. Let us choose f (A)=
∂A−σ. Under an infinitesimal transformation, Aa
μ→Aa
μ−fabcθbAc
μ+∂μθaand thus
/Delta1(A) ={/integraldisplay
Dθδ[∂Aa−σa−∂μ(fabcθbAc
μ−∂μθa)]}−1(4)
“=”{/integraldisplay
Dθδ[∂μ(fabcθbAc
μ−∂μθa)]}−1.
where the “effectively equal sign” follows since /Delta1(A) is to be multiplied later by δ[f (A)].
Let us write formally
∂μ(fabcθbAc
μ−∂μθa)=/integraldisplay
d4yKab(x,y)θb(y) (5)
thus defining the operator Kab(x,y)=∂μ(fabcAc
μ−∂μδab)δ(4)(x−y). Note that in con-
trast to electromagnetism here Kdepends on the gauge potential. The elementary result/integraltext
dθδ(Kθ) =1/K forθandKreal numbers can be generalized to/integraltext
dθδ(Kθ) =1/detK
forθa real vector and Ka nonsingular matrix. Regarding Kab(x,y)as a matrix, we ob-
tain/Delta1(A) =detK, but we know from chapter II.5 how to represent the determinant as a
functional integral over Grassmann variables: Write /Delta1(A) =/integraltext
DcDc†eiSghost(c†,c), with
Sghost(c†,c)=/integraldisplay
d4xd4yc†
a(x)Kab(x,y)cb(y)
=/integraldisplay
d4x[∂c†
a(x)∂c a(x)−∂μc†
a(x)fabcAcμ(x)cb(x)]
=/integraldisplay
d4x∂c†
a(x)Dc a(x) (6)
and with Dthe covariant derivative for the adjoint representation, to which the fields ca
andc†
abelong just like Aa
μ. The fields caandc†
aare known as ghost fields because they
violate the spin-statistics connection: Though scalar, they are treated as anticommuting.This “violation” is acceptable because they are not associated with physical particles andare introduced merely to represent /Delta1(A) in a convenient form.
This takes care of the /Delta1(A) factor in (3). As for the δ[f (A)] factor, we use the same trick
as in chapter III.4 and integrate Zoverσ
a(x) with a Gaussian weight e−(i/2ξ)/integraltext
d4xσa(x)2
so that δ[f (A)] gets replaced by e−(i/2ξ)/integraltext
d4x(∂Aa)2.
Putting it all together, we obtain
Z=/integraldisplay
DADcDc†eiS(A)−(i/2 ξ)/integraltext
d4x(∂A)2+iS ghost(c†,c)(7)
withξa gauge parameter. Comparing with the corresponding expression for an abelian
gauge theory in chapter III.4, we see that in nonabelian gauge theories we have a ghost
action Sghost in addition to the Y ang-Mills action. Thus, L0and L1are changed to
L0=−1
4(∂μAa
ν−∂νAa
μ)2−1
2ξ(∂μAa
μ)2+∂c†
a∂ca (8)
VII.1. Quantizing Yang-Mills Theory | 373
a,μ
c,λ b,νk1
k3k2a,μb,ν
d,ρc,λ
c,μ
abp(a) (b)
(c)
Figure VII.1.1
and
L1=−1
2g(∂μAa
ν−∂νAa
μ)fabcAbμAcν+1
4g2fabcfadeAbμAcνAdμAeν−∂μc†
agfabcAcμcb(x) (9)
We can now read off the propagators for the gauge boson and for the ghost field imme-
diately from (8). In particular, we see that except for the group index athe terms quadratic
in the gauge potential are exactly the same as the terms quadratic in the electromagneticgauge potential in (III.4.8). Thus, the gauge boson propagator is
(−i)
k2/bracketleftbigg
gνλ−(1−ξ)kνkλ
k2/bracketrightbigg
δab (10)
Compare with (III.4.9). From the term ∂c†
a∂cain (8) we find the ghost propagator to be
(i/k2)δab.
From L1we see that there is a cubic and a quartic interaction between the gauge
bosons, and an interaction between the gauge boson and the ghost field, as illustrated
in figure VII.1.1. The cubic and the quartic couplings can be easily read off as
gfabc[gμν(k1−k2)λ+gνλ(k2−k3)μ+gλμ(k3−k1)ν] (11)
and
−ig2[fabefcde(gμλgνρ−gμρgνλ)+fadefcbe(gμλgνρ−gμνgρλ)
+facefbde(gμνgλρ−gμρgνλ)] (12)
respectively. The coupling to the ghost field is
gfabcpμ(13)
374 | VII. Grand Unification
Obviously, we can exploit various permutation symmetries in writing these down. For
instance, in (12) the second term is obtained from the first by the interchange {c,λ}↔
{d,ρ}, and the third and fourth terms are obtained from the first and second by the
interchange {a,μ}↔{ c,λ}.
Unnatural act
In a highly symmetric theory such as Y ang-Mills, perturbating is clearly an unnatural
act as it involves brutally splitting Linto two parts: a part quadratic in the fields and
the rest. Consider, for example, an exactly soluble single particle quantum mechanicsproblem, such as the Schr ¨odinger equation with V( x)=1−(1/coshx)
2. Imagine writing
V( x)=1
2x2+W(x) and treating W(x) as a perturbation on the harmonic oscillator. You
would have a hard time reproducing the exact spectrum, but this is exactly how we brutalizeY ang-Mills theory in the perturbative approach: We took the “holistic entity” tr F
μνFμνand
split it up into the “harmonic oscillator” piece tr (∂μAν−∂νAμ)2and a “perturbation.”
If Y ang-Mills theory ever proves to be exactly soluble, the perturbative approach with its
mangling of gauge invariance is clearly not the way to do it.
Lattice gauge theory
Wilson proposed a way out: Do violence to Lorentz invariance rather than to gauge in-variance. Let us formulate Y ang-Mills theory on a hypercubic lattice in 4-dimensionalEuclidean spacetime. As the lattice spacing a→0 we expect to recover 4-dimensional
rotational invariance and (by a Wick rotation) Lorentz invariance. Wilson’s formulation,known as lattice gauge theory, is easy to understand, but the notation is a bit awkward, dueto the lack of rotational invariance. Denote the location of the lattice sites by the vector x
i.
On each link, say the one going from xito one of its nearest neighbors xj, we associate an
NbyNsimple unitary matrix Uij. Consider the square, known as a plaquette, bounded
by the four corners xi,xj,xk, and xl(with these nearest neighbors to each other.) See
figure VII.1.2. For each plaquette Pwe associate the quantity S(P)=Re trUijUjkUklUli,
constructed to be invariant under the local transformation
Uij→V†
iUijVj (14)
The symmetry is local because for each site xiwe can associate an independent Vi.
Wilson defined Y ang-Mills theory by
Z=/integraldisplay
/Pi1dUe(1/2f2)/summationtext
PS(P)(15)
where the sum is taken over all the plaquettes in the lattice. The coupling strength f
controls how wildly the unitary matrices Uij’s fluctuate. For small f, large values of S(P)
are favored, and so the Uij’s are all approximately equal to the unit matrix (up to an
irrelevant global transformation.)
VII.1. Quantizing Yang-Mills Theory | 375
xi xjxl xk
Uli UjkUkl
Uij
Figure VII.1.2
Without doing any arithmetic, we can argue by symmetry that in the continuum limit
a→0, Y ang-Mills theory as we know it must emerge: The action is manifestly invariant
under local SU(N) transformation. T o actually see this, define a field Aμ(x) withμ=
1, 2, 3, 4, permeating the 4-dimensional Euclidean space the lattice lives in, by
Uij=V†
ieiaAμ(x)Vj (16)
where x=1
2(xi+xj)(namely the midpoint of the link Uijlives on) and μis the direction
connecting xitoxj(namely ˆμ≡(xj−xi)/a is the unit vector in the μdirection.) The V’s
just reflect the gauge freedom in (14) and obviously do not enter into the plaquette actionS(P) by construction. I will let you show in an exercise that
trUijUjkUklUli=treia2Fμν+O(a3)(17)
withFμνthe Y ang-Mills field strength evaluated at the center of the plaquette. Indeed, we
could have discovered the Y ang-Mills field strength in this way. I hope that you start tosee the deep geometric significance of F
μν. Continuing the exercise you will find that the
action on each plaquette comes out to be
S(P)=Re treia2Fμν+O(a3)
=Re tr[1 +ia2Fμν−1
2a4FμνFμν+O(a5)]=tr 1−1
2a4trFμνFμν+...
(18)
and so up to an irrelevant additive constant we recover in (15) the Y ang-Mills action in the
continuum limit. Again, it is worth emphasizing that without going through any arithmeticwe could have fixed the a
4term in (18) (up to an overall constant) by dimensional analysis
and gauge invariance.1
1The sign can be easily checked against the abelian case.
376 | VII. Grand Unification
The Wilson formulation is beautiful in that none of the hand-wringing over gauge
fixing, Faddeev-Popov determinant, ghost fields, and so forth is necessary for (15) to makesense. Recalling chapter V .3 you see that (15) defines a statistical mechanics problemlike any other. Instead of integrating over some spin variables say, we integrate over thegroup SU(N) for each link. Most importantly, the lattice gauge formulation opens up
the possibility of computing the properties of a highly nontrivial quantum field theorynumerically. Lattice gauge theory is a thriving area of research. For a challenge, try toincorporate fermions into lattice gauge theory: This is a difficult and ongoing problembecause fermions and spinor fields are naturally associated with SO( 4), which does not
sit well on a lattice.
Wilson loop
Field theorists usually deal with local observables, that is, observables defined at a space-time point x, such as J
μ(x) or trFμν(x)Fμν(x), but of course we can also deal with
nonlocal observables, such as ei/contintegraltext
CdxμAμin electromagnetism, where the line integral is
evaluated over a closed curve C. The gauge invariant quantity in the exponential is equal
to the electromagnetic flux going through the surface bounded by C. (Indeed, recall chap-
ter IV .4.)
Wilson pointed out that lattice gauge theory contains a natural gauge invariant but
nonlocal observable W(C) ≡trUijUjk...UnmUmi, where the set of links connecting xi
toxjtoxket cetera and eventually to xmand back to xitraces out a loop called C. Referring
to (16) we see that W(C) , known as the Wilson loop, is the trace of a product of many
factors of eiaAμ. Thus, in the continuum limit a→0, we have evidently
W(C) ≡trPei/contintegraltext
CdxμAμ(19)
withCnow an arbitrary curve in Euclidean spacetime. Here Pdenotes path ordering,
clearly necessary since the Aμ’s associated with different segments of C, being matrices,
do not commute with each other. [Indeed, Pis defined by the lattice definition of W(C) .]
T o understand the physical meaning of the Wilson loop, Recall chapters I.4 and I.5. T o
obtain the potential energy Ebetween two oppositely charged lumps we have to compute
lim
T→∞1
Z/integraldisplay
DAeiSMaxwell (A)+i/integraltext
d4xAμJμ=e−iET
For two lumps held at a distance Rapart we plug in
Jμ(x)=ημ0{δ(3)(/vectorx)−δ(3)[/vectorx−(R,0 ,0)]}
and see that we are actually computing the expectation value /angbracketleftei(/integraltext
C1dxμAμ−/integraltext
C2dxμAμ)/angbracketrightin
a fluctuating electromagnetic field, where C1andC2denote two straight line segments
at/vectorx=(0, 0, 0 )and/vectorx=(R,0 ,0), respectively. It is convenient to imagine bringing the
two lumps together in the far future (and similarly in the far past). Then we deal instead
with the manifestly gauge invariant quantity /angbracketleftei/contintegraltext
CdxμAμ/angbracketright, where Cis the rectangle shown
VII.1. Quantizing Yang-Mills Theory | 377
T
RTime
Space
C
Figure VII.1.3
in figure VII.1.3. Note that for Tlarge log /angbracketleftei/contintegraltext
CdxμAμ/angbracketright∼−iE(R)T , which is essentially
proportional to the perimeter length of the rectangle C.
As we will discuss in chapter VII.3 and as you have undoubtedly heard, the currently
accepted theory of the strong interaction involves quarks coupled to a nonabelian Y ang-Mills gauge potential A
μ. Thus, to determine the potential energy E(R) between a quark
and an antiquark held fixed at a distance Rfrom each other we “merely” have to compute
the expectation value of the Wilson loop
/angbracketleftW(C) /angbracketright=1
Z/integraldisplay
/Pi1dUe−(1/2f2)/summationtext
PS(P)W(C) (20)
In lattice gauge theory we could compute log /angbracketleftW(C) /angbracketrightforCthe large rectangle in fig-
ure VII.1.3 numerically, and extract E(R) . (We lost the ibecause we are living in Euclidean
spacetime for the purpose of this discussion.)
Quark confinement
You have also undoubtedly heard that since free quarks have not been observed, quarks aregenerally believed to be permanently confined. In particular, it is believed that the potentialenergy between a quark and an antiquark grows linearly with separation E(R)∼σR.
One imagines a string tying the quark to the antiquark with a string tension σ. If this
conjecture is correct, then log /angbracketleftW(C) /angbracketright∼σRT should go as the area RT enclosed by C.
Wilson calls this behavior the area law, in contrast to the perimeter law characteristic offamiliar theories such as electromagnetism. T o prove the area law in Y ang-Mills theory isone of the outstanding challenges of theoretical physics.
378 | VII. Grand Unification
Exercises
VII.1.1 The gauge choice in the text preserves Lorentz invariance. It is often useful to choose a gauge that breaks
Lorentz invariance, for example, f (A)=nμAμ(x)withnsome fixed 4-vector. This class of gauge choices,
known as the axial gauge, contains various popular gauges, each of which corresponds to a particularchoice of n. For instance, in light-cone gauge, n=(1, 0, 0, 1 ), in space-cone gauge, n=(0, 1,i,0). Show
that for any given A(x) we can find a gauge transformation so that n.A
/prime(x)=0.
VII.1.2 Derive (17) and relate fto the coupling gin the continuum formulation of Y ang-Mills theory. [Hint: Use
the Baker-Campbell-Hausdorff formula
eAeB=eA+B+1
2[A,B]+1
12([A,[A,B]]+[B,[B,A]])+...
VII.1.3 Consider a lattice gauge theory in ( D+1)-dimensional space with the lattice spacing ainD-dimensional
space and bin the extra dimension. Obtain the continuum D-dimensional field theory in the limit a→0
withbkept fixed.
VII.1.4 Study in (2) the alternative limit b→0 with akept fixed so that you obtain a theory on a spatial lattice
but with continuous time.
VII.1.5 Show that for lattice gauge theory the Wilson area law holds in the limit of strong coupling. [Hint: Expand
(20) in powers of f−2.]
VII.2 Electroweak Unification
The scourge of massless spin 1 particles
With the benefit of hindsight, we now know that Nature likes Y ang-Mills theory. In the
late 1960s and early 1970s, the electromagnetic and weak interactions were unified intoan electroweak interaction, described by a nonabelian gauge theory based on the groupSU( 2)⊗U(1). Somewhat later, in the early 1970s, it was realized that the strong interaction
can be described by a nonabelian gauge theory based on the group SU( 3). Nature literally
consists of a web of interacting Y ang-Mills fields.
But when the theory was first proposed in 1954, it seemed to be totally inconsistent with
observations as they were interpreted at that time. As Y ang and Mills themselves pointedout in their paper, the theory contains massless spin 1 particles, which were certainly notknown experimentally. Thus, except for interest on the part of a few theorists (Schwinger,Glashow, Bludman, and others) who found the mathematical structure elegantly attractiveand felt that nonabelian gauge theory must somehow be relevant for the weak interaction,the theory gradually sank into oblivion and was not part of the standard graduate curricu-lum in particle physics in the 1960s.
Again with the benefit of hindsight, it would seem that there are only two logical
solutions to the difficulty that experimentalists do not see any massless spin 1 particlesexcept for the photon: (1) the Y ang-Mills particles somehow acquire mass, or (2) the Y ang-Mills particles are in fact massless but are somehow not observed. We now know that thefirst possibility was realized in the electroweak interaction and the second in the stronginteraction.
Constructing the electroweak theory
We now discuss electroweak unification. It is perhaps pedagogically clearest to motivate
how we would go about constructing such a theory. As I have said before, this is not a
380 | VII. Grand Unification
textbook on particle physics and I necessarily will have to keep the discussion of particle
physics to the bare minimum. I gave you a brief introduction to the structure of the weakinteraction in chapter IV .2. The other salient fact is that weak interaction violates parity,as mentioned in chapter II.1. In particular, the left handed electron field e
Land the right
handed electron field eR, which transform into each other under parity, enter into the weak
interaction quite differently.
Let us start with the weak decay of the muon, μ−→e−+¯ν+ν/prime, with νandν/primethe
electron neutrino and muon neutrino, respectively. The relevant term in the Lagrangianis¯ν
/prime
LγμμL¯eLγμνL, with the left hand electron field eL, the electron neutrino field (which
is left handed) νL, and so forth. The field μLannihilates a muon, the field ¯eLcreates an
electron, and so on. (Henceforth, we will suppress the word field.) As you probably know,the elementary constituents of matter form three families, with the first family consistingofν,e, and the up uand down dquarks, the second of ν
/prime,μ, and the charm cand strange
squarks, and so on. For our purposes here, we will restrict our attention to the first family.
Thus, we start with ¯νLγμeL¯eLγμνL.
As I remarked in chapter III.2, a Fermi interaction of this type can be generated by the
exchange of an intermediate vector boson W+with the coupling W+
μ¯νLγμeL+W−
μ¯eLγμνL.
The idea is then to consider an SU( 2)gauge theory with a triplet of gauge bosons denoted
byWa
μ, with a=1, 2, 3. Put νLandeLinto the doublet representation and the right handed
electron field eRinto a singlet representation, thus
ψL≡/parenleftBiggν
e/parenrightBigg
L,eR (1)
(The notation is such that the upper component of ψLisνLand the lower component is
eL.)
The fields νLandeL, but not eR, listen to the gauge bosons Wa
μ. Indeed, according to
(IV .5.21) the Lagrangian contains
Wa
μ¯ψLτaγμψL=(W1−i2
μ¯ψL1
2τ1+i2γμψL+h.c.)+W3
μ¯ψLτ3γμψL
where W1−i2
μ≡W1
μ−iW2
μand so forth. We recognize τ1+i2≡τ1+iτ2as the raising
operator and the first two terms as (W1−i2
μ¯νLγμeL+h.c.), precisely what we want. By
design, the exchange of W±
μgenerates the desired term ¯νLγμeL¯eLγμνL.
We need more room
We would hope that the boson W3we were forced to introduce would turn out to be the pho-
ton so that electromagnetism is included. But alas, W3couples to the current ¯ψLτ3γμψL=
(¯νLγμνL−¯eLγμeL), not the electromagnetic current −(¯eLγμeL+¯eRγμeR). Oops!
Another problem lurks. T o generate a mass term for the electron, we need a doublet
Higgs field ϕ≡/parenleftbigϕ+
ϕ0/parenrightbig
in order to construct the SU( 2)invariant term f¯ψLϕeRin the
Lagrangian so that when ϕacquires the vacuum expectation value/parenleftbig0
v/parenrightbig
we will have
VII.2. Electroweak Unification | 381
f¯ψLϕeR→f(¯ν,¯e)L/parenleftBigg0
v/parenrightBigg
eR=fv¯eLeR (2)
But none of the SU( 2)transformations leaves/parenleftbig0
v/parenrightbig
invariant: The vacuum expectation value
ofϕspontaneously breaks the entire SU( 2)symmetry, leaving all three Wbosons massive.
There is no room for the photon in this failed theory. Aagh!
We need more room. Remarkably, we can avoid both the oops and the aagh by extending
the gauge symmetry to SU( 2)⊗U(1). Denoting the generator of U(1)by1
2Y(called the
hypercharge) and the associated gauge potential by Bμ[and their counterparts TaandWa
μ
forSU( 2)] we have the covariant derivative Dμ=∂μ−igWa
μTa−ig/primeBμY
2. With four gauge
bosons, we dare to hope that one of them might turn out to be the photon.
The gauge potentials are normalized by the corresponding kinetic energy terms,
L=−1
4(Bμν)2−1
4(Wa
μν)2+...with the abelian Bμν=∂μBν−∂νBμand nonabelian field
strength Wa
μν=∂μWa
ν−∂νWa
μ+εabcWb
μWc
ν. The generators Taare of course normalized
by the commutation relations that define SU( 2). In contrast, there is no commutation
relation in the abelian algebra U(1)to fix the normalization of the generator1
2Y. Until this
is fixed, the normalization of the U(1)gauge coupling g/primeis not fixed.
How do we fix the normalization of the generator1
2Y? By construction, we want spon-
taneous symmetry breaking to leave a linear combination of T3and1
2Yinvariant, to be
identified as the generator the massless photon couples to, namely the charge operator Q.
Thus, we write
Q=T3+1
2Y (3)
Once we know T3and1
2Yof any field, this equation tells us its charge. For example,
Q(νL)=1
2+1
2Y(νL)andQ(eL)=−1
2+1
2Y(eL). In particular, we see that the coefficient
ofT3in (3) must be 1 since the charges of νLandeLdiffer by 1. The relation (3) fixes the
normalization of1
2Y.
Determining the hypercharge
The next step is to determine the hypercharge of various multiplets in the theory, which in
turn determines how Bμcouples to these multiplets. Consider ψL.F o reLto have charge
−1, the doublet ψLmust have1
2Y=−1
2. In contrast, the field eRhas1
2Y=− 1 since T3=0
oneR.
Given the hypercharge of ψLandeRwe see that the invariance of the term f¯ψLϕeR
under SU( 2)⊗U(1)forces the Higgs field ϕto have1
2Y=+1
2. Thus, according to (3)
the upper component of ϕhas electric charge Q=+1
2+1
2=+ 1 and the lower component
Q=−1
2+1
2=0. Thus, we write ϕ=/parenleftbigϕ+
ϕ0/parenrightbig
. Recall that ϕhas the vacuum expectation value/parenleftbig0
v/parenrightbig
. The fact that the electrically neutral field ϕ0acquires a vacuum expectation value but
the charged field ϕ+does not provide a consistency check.
382 | VII. Grand Unification
The theory works itself out
Now that the couplings of the gauge bosons to the various fields, in particular, the Higgs
field, are determined, we can easily work out the mass spectrum of the gauge bosons, asindeed, let me remind you, you have already done in exercise IV .6.3!
Upon spontaneous symmetry breaking ϕ→(1/√
2)/parenleftbig0
v/parenrightbig
(the normalization is conven-
tional): We simply plug in
L=(Dμϕ)†(Dμϕ)→g2v2
4W+
μW−μ+v2
8(gW3
μ−g/primeBμ)2(4)
I trust that this is what you got! Thus, the linear combination gW3
μ−g/primeBμbecomes massive
while the orthogonal combination remains massless and is identified with the photon. Itis clearly convenient to define the angle θby tan θ=g
/prime/g. Then,
Zμ=cosθW3
μ−sinθBμ (5)
describes a massive gauge boson known as the Zboson, while the electromagnetic po-
tential is given by Aμ=sinθW3
μ+cosθBμ. Combine (4) and (5) and verify that the mass
squared of the Zboson is M2
Z=v2(g2+g/prime2)/4, and thus by elementary trigonometry ob-
tain the relation
MW=MZcosθ (6)
The exchange of the Wboson generates the Fermi weak interaction
L=−g2
2M2
W¯νLγμeL¯eLγμνL=−4G√
2¯νLγμeL¯eLγμνL
where the second equality merely gives the historical definition of the Fermi coupling G.
Thus,
G√
2=g2
8M2
W(7)
Next, we write the relevant piece of the covariant derivative
gW3
μT3+g/primeBμY
2=g(cosθZμ+sinθAμ)T3+g/prime(−sinθZμ+cosθAμ)Y
2
in terms of the physically observed ZandA. The coefficient of Aμworks out to be
gsinθT3+g/primecosθ(Y/ 2)=gsinθ(T3+Y/2); the fact that the combination Q=T3+
Y/2 emerges provides a nice check on the formalism. Furthermore, we obtain
e=gsinθ (8)
Meanwhile, it is convenient to write gcosθT3−g/primesinθ(Y/ 2), the coefficient of Zμin
the covariant derivative, in terms of the physically familiar electric charge Qrather than
the theoretical hypercharge Y: Thus,
gcosθT3−g/primesinθ(Q−T3)=g
cosθ(T3−sin2θQ)
VII.2. Electroweak Unification | 383
In other words, we have determined the coupling of the Zboson to an arbitrary fermion
field/Psi1in the theory:
L=g
cosθZμ¯/Psi1γμ(T3−sin2θQ)/Psi1 (9)
For example, using (9) we can immediately write the coupling of Zto leptons:
L=g
cosθZμ[1
2(¯νLγμνL−¯eLγμeL)+sin2θ¯eγμe] (10)
Including quarks
How to include the hadrons is now almost self evident. Given that only left handed
fields participate in the weak interaction, we put the quarks of the first generation intoSU( 2)⊗U(1)multiplets as follows:
qα
L≡/parenleftBigguα
dα/parenrightBigg
L,uα
R,dα
R(11)
where α=1, 2, 3 denotes the color index, which I will discuss in the next chapter. The
right handed quarks uα
Randdα
Rare put into singlets so that they do not hear the weak
bosons Wa. Recall that the up quark uand the down quark dhave electric charges2
3and
−1
3respectively. Referring to (3) we see1
2Y=1
6,2
3, and−1
3forqα
L,uα
R, anddα
R, respectively.
From (9) we can immediately read off the coupling of the Zboson to the quarks:
L=g
cosθZμ[1
2(¯uLγμuL−¯dLγμdL)−sin2θJμ
em] (12)
Finally, I leave it to you to verify that of the four degrees of freedom contained in ϕ
(since ϕ+andϕ0are complex) three are eaten by the WandZbosons, leaving one physical
degree of freedom Hcorresponding to the elusive Higgs particle that experimenters are
still searching for as of this writing.
The neutral current
By virtue of its elegantly economical gauge group structure, this SU( 2)⊗U(1)electroweak
theory of Glashow, Salam, and Weinberg ushered in the last great predictive era of theo-
retical particle physics. Writing (10) and (12) as
L=g
cosθZμ(Jμ
leptons+Jμ
quarks)
and using (6) we see that Zboson exchange generates a hitherto unknown neutral current
interaction
Lneutral current =−g2
2M2
W(Jleptons +Jquarks)μ(Jleptons +Jquarks)μ
between leptons and quarks. By studying various processes described by Lneutral current we
can determine the weak angle θ. Once θis determined, we can predict gfrom (8). Once g
384 | VII. Grand Unification
is determined, we can predict MWfrom (7). Once MWis determined, we can predict MZ
from (6).
Concluding remarks
As I mentioned, there are three families of leptons and quarks in Nature, consisting
of(νe,e,u,d),(νμ,μ,c,s), and (ντ,τ,t,b). The appearance of this repetitive family
structure, about which the SU( 2)⊗U(1)theory has nothing to say, represents one of
the great unsolved puzzles of particle physics. The three families, with the appropriaterotation angles between them, are simply incorporated into the theory by repeating whatwe wrote above.
A more logical approach than the one given here would be to start with an SU( 2)⊗U(1)
theory with a doublet Higgs field with some hypercharge, and to say, “Behold, uponspontaneous symmetry breaking, one linear combination of generators remains unbrokenwith a corresponding massless gauge field.” I think that our quasi-historical approach isclearer.
As I have mentioned on several occasions, Fermi’s theory of the weak interaction is
nonrenormalizable. In 1999, ’t Hooft and Veltman were awarded the Nobel Prize forshowing that the SU( 2)⊗U(1)electroweak theory is renormalizable, thus paving the way
for the triumph of nonabelian gauge theories in describing the strong, electromagnetic,and weak interactions. I cannot go into the details of their proof here, but I would liketo mention that the key is to start with the nonabelian analog of the unitary gauge (recallchapter IV .6) and proceed to the R
ξgauge. At large momenta, the massive gauge boson
propagators go as ∼(kμkν/k2)in the unitary gauge, but as ∼(1/k2)in the Rξgauge. The
theory is then renormalizable by power counting.
Exercises
VII.2.1 Unfortunately, the mass of the elusive Higgs particle Hdepends on the parameters in the double well
potential V=−μ2ϕ†ϕ+λ(ϕ†ϕ)2responsible for the spontaneous symmetry breaking. Assuming that
His massive enough to decay into W++W−andZ+Z, determine the rates for Hto decay into various
modes.
VII.2.2 Show that it is possible to stay with the SU( 2)gauge group and to identify W3as the photon A, but at
the cost of inventing some experimentally unobserved lepton fields. This theory does not describe ourworld: For one thing, it is essentially impossible to incorporate the quarks. Show this! [Hint: We have toput the leptons into a triplet of SU( 2)instead of a doublet.]
VII.3 Quantum Chromodynamics
Quarks
Quarks come in six flavors, known as up, down, strange, charm, bottom, and top, denoted
byu,d,s,c,b, andt. The proton, for example, is made of two up quarks and a down quark
∼(uud), while the neutral pion corresponds to ∼(u¯u−d¯d)/√
2. Please consult any text
on particle physics for details.
By the late 1960s the notion of quarks was gaining wide acceptance, but two separate
lines of evidence indicated that a crucial element was missing. In studying how hadronsare made of quarks, people realized that the wave function of the quarks in a nucleon doesnot come out to be antisymmetric under the interchange of any pair of quarks, as requiredby the Pauli exclusion principle. At around the same time, it was realized that in the idealworld we used to derive the Goldberger-T reiman relation and in which the pion is masslesswe can calculate the decay rate for the process π
0→γ+γ, as mentioned in chapter IV .7.
Puzzlingly enough, the calculated rate came out smaller than the observed rate by a factor
of 9=32.
Both puzzles could be resolved in one stroke by having quarks carry a hitherto unknown
internal degree of freedom that Gell-Mann called color. For any specified flavor, a quarkcomes in one of three colors. Thus, the up quark can be red, blue, or yellow. In a nucleon,the wave function of the three quarks will then contain a factor referring to color, besidesthe factors referring to orbital motion, spin, and so on. We merely have to make the colorpart of the wave function antisymmetric; in fact, we simply take it to be ε
αβγ, where α,β,
andγdenote the colors carried by the three quarks. With quarks in three colors, we have to
multiply the amplitude for π0decay by a factor of 3, thus neatly resolving the discrepancy
between theory and experiment.
386 | VII. Grand Unification
Asymptotic freedom
As I mentioned in chapter VI.6, the essential clue came from studying deep inelastic scat-
tering of electrons off nucleons. Experimentalists made the intriguing discovery that whenhit hard the quarks in the nucleons act as if they hardly interact with each other, in otherwords, as if they are free. On the other hand, since quarks are never seen as isolated enti-ties, they appear to be tightly bound to each other within the nucleon. As I have explained,this puzzling and apparently contradictory behavior of the quarks can be understood if thestrong interaction coupling flows to zero in the large momentum (ultraviolet) limit and toinfinity or at least to some large value in the small momentum (infrared) limit. A numberof theorists proposed searching for theories whose couplings would flow to zero in theultraviolet limit, now known as asymptotic free theories. Eventually, Gross, Wilczek, andPolitzer discovered that Y ang-Mills theory is asymptotically free.
This result dovetails perfectly with the realization that quarks carry color. The nonabelian
gauge transformation would take a quark of one color into a quark of another color. Thus, towrite down the theory of the strong interaction we simply take the result of exercise IV .5.6,
L=−1
4g2Fa
μνFaμν+¯q(iγμDμ−m)q (1)
with the covariant derivative Dμ=∂μ−iAμ. The gauge group is SU( 3)with the quark field
qin the fundamental representation. In other words, the gauge fields Aμ=Aa
μTa, where
Ta(a=1 ,...,8 )are traceless hermitean 3 by 3 matrices. Explicitly, ( Aμq)α=Aa
μ(Ta)αβqβ,
where α,β=1, 2, 3. The theory is known as quantum chromodynamics, or QCD for short,
and the nonabelian gauge bosons are known as gluons. T o incorporate flavor, we simplywrite/summationtext
f
j=1¯qj(iγμDμ−mj)qjfor the second term in (1), where the index jgoes over the
fflavors. Note that quarks of different flavors have different masses.
Infrared slavery
The flip side of asymptotic freedom is infrared slavery. We cannot follow the renormaliza-
tion group flow all the way down to the low momentum scale characteristic of the quarksbound inside hadrons since the coupling gbecomes ever stronger and our perturbative cal-
culation of β(g) is no longer adequate. Nevertheless, it is plausible although never proven
thatggoes to infinity and that the gluons keep the quarks and themselves in permanent
confinement. The Wilson loop introduced in chapter VII.1 provides the order parameterfor confinement.
In elementary physics forces decrease with the separation between interacting objects,
so permanent confinement is a rather bizarre concept. Are there any other instances ofpermanent confinement?
Consider a magnetic monopole in a superconductor. We get to combine what we learned
in chapters IV .4 and V .4 (and even VI.2)! A quantized amount of magnetic flux comes outof the monopole, but according to the Meissner effect a superconductor expels magneticflux. Thus, a single magnetic monopole cannot live inside a superconductor.
VII.3. Quantum Chromodynamics | 387
M M
R
Figure VII.3.1
Now consider an antimonopole a distance Raway (figure VII.3.1). The magnetic flux
coming out of the monopole can go into the antimonopole, forming a tube connectingthe monopole and the antimonopole and obliging the superconductor to give up being asuperconductor in the region of the flux tube. In the language of chapter V .4, it is no longerenergetically favorable for the field or order parameter ϕto be constant everywhere; instead
it vanishes in the region of the flux tube. The energy cost of this arrangement evidentlygrows as R(consistent with Wilson’s area law).
In other words, an experimentalist living inside a superconductor would find that
it costs more and more energy to pull a monopole and an antimonopole apart. Thisconfinement of monopoles inside a superconductor is often taken to be a model of the yet-to-be-proven confinement of quarks. Invoking electromagnetic duality we can imaginea magnetic superconductor in contrast to the usual electric superconductor. Inside amagnetic superconductor, electric charges would be permanently confined. Our universemay be likened to a color magnetic superconductor in which quarks (the analog of electriccharges) are confined.
On distance scales large compared to the radius of the color flux tube connecting a quark
to an antiquark, the tube can be thought of as a string. Historically, that was how stringtheory originated. The challenge, boys and girls, is to prove that the ground state or vacuumof (1) is a color magnetic superconductor.
Symmetries of the strong interaction
Now that we have a theory of the strong interaction, we can understand the origin of thesymmetries of the strong interaction, namely the isospin symmetry of Heisenberg and thechiral symmetry that when spontaneously broken leads to the appearance of the pion as aNambu-Goldstone boson (as discussed in chapters IV .2 and VI.4).
Consider a world with two flavors, which is all that is relevant for a discussion of the pion.
Introduce the notation u≡q
1,d≡q2, andq=/parenleftbigu
d/parenrightbig
so that we can write the Lagrangian as
L=−1
4g2Fa
μνFaμν+¯q(iγμDμ−m)q
with
m=/parenleftBiggmu 0
0md/parenrightBigg
388 | VII. Grand Unification
where muandmdare the masses of the up and down quarks, respectively. If mu=md,
the Lagrangian is invariant under q→eiθ.τq, corresponding to Heisenberg’s isospin
symmetry.
In the limit in which muandmdvanish, the Lagrangian is invariant under q→eiϕ.τγ5q,
known as the chiral SU( 2)symmetry, chiral because the right handed quarks qRand the
left handed quarks qLtransform differently. T o the extent that muandmdare both much
smaller than the energy scale of the strong interaction, chiral SU( 2)is an approximate
symmetry.
The pion is the Nambu-Goldstone boson associated with the spontaneous breaking
of the chiral SU( 2). Indeed, this is an example of dynamical symmetry breaking since
there is no elementary scalar field around to acquire a vacuum expectation. Instead, thestrong interaction dynamics is supposed to drive the composite scalar fields ¯uuand¯ddto
“condense into the vacuum” so that /angbracketleft0|¯uu|0/angbracketright=/angbracketleft 0|¯dd|0/angbracketrightbecome nonvanishing, where the
equality between the two vacuum expectation values ensures that Heisenberg’s isospin isnot spontaneously broken, an experimental fact since there are no corresponding Nambu-Goldstone bosons. In terms of the doublet field q, the QCD vacuum is supposed to be such
that/angbracketleft0|¯qq|0/angbracketright /negationslash= 0 while /angbracketleft0|¯q/vectorτq|0/angbracketright=0.
Renormalization group flow
The renormalization group flow of the QCD coupling is governed by
dg
dt=β(g)=−11
3T2(G)g3
16π2(2)
with the all-crucial minus sign. Here
T2(G)δab=facdfbcd(3)
I will not go through the calculation of β(g) here, but having mastered chapters VI.8
and VII.1 you should feel that you can do it if you want to.1At the very least, you should
understand the factor g3andT2(G) by drawing the relevant Feynman diagrams.
When fermions are included,
dg
dt=β(g)=/bracketleftBig
−11
3T2(G)+4
3T2(F)/bracketrightBigg3
16π2(4)
where
T2(F)δab=tr[Ta(F)Tb(F)] (5)
I do expect you to derive (4) given (2). For SU(N) T2(F)=1
2for each fermion in the
fundamental representation. Note that asymptotic freedom is lost when there are too many
fermions.
1For a detailed calculation, see, e.g., S. Weinberg, The Quantum Theory of Fields , Vol. 2, sec. 18.7.
VII.3. Quantum Chromodynamics | 389
γ
e–e+
Figure VII.3.2
You already solved an equation like (4) in exercise VI.8.1. Let us define, in analogy to
quantum electrodynamics, αS(μ)≡g(μ)2/4π, the strong coupling at the momentum scale
μ. From (4) we obtain2
αS(Q)=αS(μ)
1+(1/4π)(11−2
3nf)αS(μ) log(Q2/μ2)(6)
showing explicitly that αS(Q)→0 logarithmically as Q→∞ .
Electron-positron annihilation
I have space to show you only one physical application. Experimentalists have measured
the cross section σofe+e−annihilation into hadrons as a function of the total center-of-
mass energy E. The amplitude is shown in figure VII.3.2. T o calculate the cross section in
terms of the amplitude, we have to go through what some people call “boring kinematics,”such as normalizing everything correctly, dividing by the flux of the two beams, and soforth (see the appendix to chapter II.6). For the good of your soul, you should certainly gothrough this type of calculation at least once. Believe me, I did it more times than I careto remember. But happily, as I will now show you, we can avoid most of this grunge labor.First, consider the ratio
R(E)≡σ(e+e−→hadrons)
σ(e+e−→μ+μ−)
The kinematic stuff cancels out. In figure VII.3.2 the half of the diagram involving the
electron positron lines and the photon propagator also appears in the Feynman diagram
e+e−→μ+μ−(figure VII.3.3) and so cancels out in R(E) . The blob in figure VII.3.2, which
hides all the complexity of the strong interaction, is given by /angbracketleft0|Jμ(0)|h/angbracketright, where Jμis the
2For the accumulated experimental evidence on αS(Q) , see figure 14.3 in F. Wilczek, in: V . Fitch et al., eds.,
Critical Problems in Physics , p. 281.
390 | VII. Grand Unification
γ
e−e+μ−μ+
Figure VII.3.3
electromagnetic current and the state |h/angbracketrightcan contain any number of hadrons. T o obtain
the cross section we have to square the amplitude, include a δ-function for momentum
conservation, and sum over all |h/angbracketright, thus arriving at
/summationdisplay
h(2π)4δ4(ph−pe+−pe−)/angbracketleft0|Jμ(0)|h/angbracketright/angbracketlefth| Jν(0)|0/angbracketright (7)
[withq≡pe++pe−=(E,/vector0)]. This quantity can be written as
/integraldisplay
d4xeiqx/angbracketleft0|Jμ(x)Jν(0)|0/angbracketright=/integraldisplay
d4xeiqx/angbracketleft0|[Jμ(x),Jν(0)]|0/angbracketright
=2I m(i/integraldisplay
d4xeiqx/angbracketleft0|TJμ(x)Jν(0)|0/angbracketright)
(The first equality follows from E> 0 and the second was explained in chapter III.8.)
T o determine this quantity, we would have to calculate an infinite number of Feynmandiagrams involving lots of quarks and gluons. A typical diagram is shown in figure VII.3.4.Completely hopeless!
This is where asymptotic freedom rides to the rescue! From chapter VI.7 you learned
that for a process at energy Ethe appropriate coupling strength to use is g(E) . But as we
crank up E,g(E) gets smaller and smaller. Thus diagrams such as figure VII.3.4 involving
many powers of g(E) all fall away, leaving us with the diagrams with no power of g(E)
(fig. VII.3.5a) and two powers of g(E) (figs. VII.3.5b,c,d). No calculation is necessary to
obtain the leading term in R(E) , since the diagram in figure VII.3.5a is the same one
that enters into e
+e−→μ+μ−: We merely replace the quark propagator by the muon
propagator (quark and muon masses are negligible compared to E). At high energy, the
quarks are free and R(E) merely counts the square of the charge Qaof the various quarks
contributing at that energy. We predict
R(E) −→
E→∞3/summationdisplay
aQ2
a(8)
The factor of 3 accounts for color.
VII.3. Quantum Chromodynamics | 391
γ
γ
Figure VII.3.4
Not only does QCD turn itself off at high energies, it tells us how fast it is turning itself
off. Thus, we can determine how the limit in (8) is approached:
R(E)=/parenleftBigg
3/summationdisplay
aQ2
a/parenrightBigg/parenleftBigg
1+C2
(11−2
3nf)log(E/μ)+.../parenrightBigg
(9)
I will leave it to you to calculate C.
Dreams of exact solubility
An analytic solution of quantum chromodynamics is something of a “Holy Grail” for
field theorists (a grail that now carries a prize of one million dollars: see www.ams.org/claymath/). Many field theorists have dreamed that at least “pure” QCD, that is QCDwithout quarks, might be exactly soluble. After all, if any 4-dimensional quantum field
(a) (b) (c) (d)γ
γ
Figure VII.3.5
392 | VII. Grand Unification
1
QCD μg(μ)2
4π
Figure VII.3.6
theory turns out to be exactly soluble, pure Y ang-Mills, with all its fabulous symmetries,
is the most likely possibility. (Perhaps an even more likely candidate for solubility issupersymmetric Y ang-Mills theory. We will touch on supersymmetry in chapter VIII.4.)
Let me be specific about what it means to solve QCD. Consider a world with only up
and down quarks with m
uandmdboth set equal to zero, namely a world described by
L=−1
4g2Fa
μνFaμν+¯qiγμDμq (10)
The goal would be to calculate something like the ratio of the mass of the ρmeson mρto
the mass of the proton mP.
T o make progress, theoretical physicists typically need to have a small parameter to
expand in, but in trying to solve (10) we are confronted with the immediate difficultythat there is no such parameter. You might think that gis a parameter, but you would
be mistaken. The renormalization group analysis taught us that g(μ) is a function of the
energy scale μat which it is measured. Thus, there is no particular dimensionless number
we can point to and say that it measures the strength of QCD. Instead, the best we can dois to point to the value of μat which (g(μ)
2/4π)becomes of order 1. This is the energy,
known as /Lambda1QCD , at which the strong interaction becomes strong as we come down from
high energy (fig. VII.3.6). But /Lambda1QCD merely sets the scale against which other quantities
are to be measured. In other words, if you manage to calculate mPit better come out
proportional to /Lambda1QCD since/Lambda1QCD is the only quantity with dimension of mass around.
Similarly for mρ. Put in precise terms, if you publish a paper with a formula giving mρ/mP
in terms of pure numbers such as 2 and π, the field theory community will hail you as a
conquering hero who has solved QCD exactly.
The apparent trade of a dimensionless coupling gfor a dimensional mass scale /Lambda1QCD
is known as dimensional transmutation, of which we will see another example in the next
chapter.
VII.3. Quantum Chromodynamics | 393
Exercises
VII.3.1 Calculate Cin (9). [Hint: If you need help, consult T . Appelquist and H. Georgi, Phys. Rev. D8: 4000, 1973;
and A. Zee, Phys. Rev. D8: 4038, 1973.]
VII.3.2 Calculate (2).
VII.4 LargeNExpansion
Inventing an expansion parameter
Quantum chromodynamics is a zero-parameter theory, so it is difficult to give even a first
approximation. In desperation, field theorists invented a parameter in which to expandQCD. Suppose instead of three colors we have Ncolors. ’t Hooft
1noticed that as N→∞
remarkable simplifications occur. The idea is that if we can calculate mρ/mP, for example,
in the large Nlimit the result may be close to the actual value. People sometimes joke that
particle physicists regard 3 as a large number, but actually the correction to the large N
limit is typically of order 1 /N2, about 10% in the real world. Particle physicists would be
more than happy to be able to calculate hadron masses to this degree of accuracy.
As with spontaneous symmetry breaking and a number of other important concepts, the
largeNexpansion came out of condensed matter physics but nowadays is used routinely
in all sorts of contexts. For example, people have tried a large Napproach to solve high-
temperature superconductivity and to fold RNA.2
Scaling the QCD coupling
So, let the color group be U(N) and write
L=−Na
2g2trFμνFμν+¯ψ[i(/negationslash∂−i/negationslashA)−m]ψ (1)
Note that we have replaced g2byg2/Na. For finite Nthis change has no essential signifi-
cance. The point is to choose the power aso that interesting simplifications occur in the
limitN→∞ withg2held fixed. The cubic and quartic interaction vertices of the gluons
1G. ’t Hooft, Under the Spell of the Gauge Principle , p. 378.
2M. Bon, G. Vernizzi, H. Orland, and A. Zee, “T opological classification of RNA structures,” J. Mol. Biol.
379:900, 2008.
VII.4. Large NExpansion | 395
γγ γγ
(a) (b)
Figure VII.4.1
are proportional to Na. On the other hand, since the gluon propagator goes as the inverse
of the quadratic terms in L, it is proportional to 1 /Na. The coupling of the gluon to the
quark does not depend on N.
To fi x a, let us focus on a specific application, the calculation of σ(e+e−→hadrons)
discussed in the last chapter. Suppose we want to calculate this cross section at lowenergies. Consider the two-gluon exchange diagrams shown in figures VII.4.1a and b.The two diagrams are of order g
4and we would have to calculate both. Note that 1b is
nonplanar: Since one gluon crosses over the other, the diagram cannot be drawn on theplane if we insist that lines cannot go through each other.
Now the double-line formalism introduced in chapter IV .5 shines. In this formalism
the diagrams figure VII.4.1a and b are redrawn as in figure VII.4.2a and b. The two gluonpropagators common to both diagrams give a factor 1 /N
2a. Now comes the punchline.
We sum over three independent color indices in 2a, thus getting a factor N3. Grab some
crayons and try to color each line in 2a with a different color: you will need three crayons.In contrast, we sum over only one independent index in 2b, getting only a factor of N.I n
other words, 2a dominates 2b by a factor N
2. In the large Nlimit we can throw 2b away.
Clearly, the rule is to associate one factor of Nwith each loop. Thus, the lowest order
diagram, shown in 2c, with Ndifferent colors circulating in it, scales as N; 2a scales as
N3/N2a. We want 2a and 2c to scale in the same way and thus we choose a=1.
(a) (b)
(c) (d)
Figure VII.4.2
396 | VII. Grand Unification
By drawing more diagrams [e.g., 2d scales as N(1/N4)N4, with the three factors coming
from the quartic coupling, the propagators, and the sum over colors, respectively], you canconvince yourself that planar diagrams dominate in the large Nlimit, all scaling as N.F o r
a challenge, try to prove it. Evidently, there is a topological flavor to all this.
The reduction to planar diagrams is a vast simplification but there are still an infinite
number of diagrams. At this stage in our mastery of field theory, we still can’t solve largeNQCD. (As I started writing this book, there were tantalizing clues, based on insight and
techniques developed in string theory, that a solution of large NQCD might be within
sight. As I now go through the final revision, that hope has faded.)
The double-line formalism has a natural interpretation. Group theoretically, the matrix
gauge potential A
i
jtransforms just like ¯qiqj(but assuredly we are not saying that the gluon
is a quark-antiquark bound state) and the two lines may be thought of as describing a quarkand an antiquark propagating along, with the arrows showing the direction in which coloris flowing.
Random matrix theory
There is a much simpler theory, structurally similar to large NQCD, that actually can be
solved. I am referring to random matrix theory.
Exaggerating a bit, we can say that quantum mechanics consists of writing down a
matrix known as the Hamiltonian and then finding its eigenvalues and eigenvectors. In theearly 1950s, when confronted with the problem of studying the properties of complicatedatomic nuclei, Eugene Wigner proposed that instead of solving the true Hamiltonian insome dubious approximation we might generate large matrices randomly and study thedistribution of the eigenvalues—a sort of statistical quantum mechanics. Random matrixtheory has since become a rich and flourishing subject, with an enormous and growingliterature and applications to numerous areas of theoretical physics and even to puremathematics (such as operator algebra and number theory.)
3It has obvious applications
to disordered condensed matter systems and less obvious applications to random surfacesand hence even to string theory. Here I will content myself with showing how ’t Hooft’sobservation about planar diagrams works in the context of random matrix theory.
Let us generate NbyNhermitean matrices ϕrandomly according to the probability
P(ϕ)=1
Ze−N trV( ϕ)(2)
withV( ϕ) a polynomial in ϕ. For example, let V( ϕ)=1
2m2ϕ2+gϕ4. The normalization/integraltext
dϕP(ϕ) =1 fixes
Z=/integraldisplay
dϕe−N trV( ϕ)(3)
The limit N→∞ is always understood.
3For a glimpse of the mathematical literature, see D. Voiculescu, ed., Free Probability Theory .
VII.4. Large NExpansion | 397
As in chapter VI.7 we are interested in ρ(E) , the density of eigenvalues of ϕ. T o make
sure that you understand what is actually meant, let me describe what we would do werewe to evaluate ρ(E) numerically. For some large integer N, we would ask the computer to
generate a hermitean matrix ϕwith the probability P(ϕ) and then to solve the eigenvalue
equation ϕv=Ev. After this procedure had been repeated many times, the computer could
plot the distribution of eigenvalues in a histogram that eventually approaches a smoothcurve, called the density of eigenvalues ρ(E) .
We already developed the formalism to compute ρ(E) in (VI.7.1): Compute the real
analytic function G(z)≡/angbracketleft(1/N) tr[1/(z−ϕ)]/angbracketrightandρ(E)=−(1/π) lim
ε→0ImG(E+iε). The
average /angbracketleft.../angbracketrightis taken with the probability P(ϕ) :
/angbracketleftO(ϕ) /angbracketright=1
Z/integraldisplay
Dϕe−N trV( ϕ)O(ϕ)
You see that my choice of notation, ϕfor the matrix and V( ϕ)=1
2m2ϕ2+gϕ4as an
example, is meant to be provocative. The evaluation of Zis just like the evaluation of a path
integral, but for an action S(ϕ)=NtrV( ϕ) that does not involve/integraltext
ddx. Random matrix
theory can be thought of as a quantum field theory in (0+0)-dimensional spacetime!
Various field theoretic methods, such as Feynman diagrams, can all be applied to
random matrix theory. But life is sweet in (0+0)-dimensional spacetime: There is no
space, no time, no energy, and no momentum and hence no integral to do in evaluatingFeynman diagrams.
The Wigner semicircle law
Let us see how this works for the simple case V( ϕ)=1
2m2ϕ2(we can always absorb minto
ϕbut we won’t). Instead of G(z), it is slightly easier to calculate
Gi
j(z)≡/angbracketleftBigg/parenleftbigg1
z−ϕ/parenrightbiggi
j/angbracketrightBigg
=δi
jG(z)
The last equality follows from invariance under unitary transformations:
P(ϕ)=P(U†ϕU) (4)
Expand
Gi
j(z)=∞/summationdisplay
n=01
z2n+1/angbracketleft(ϕ2n)i
j/angbracketright (5)
Do the Gaussian integral
1
Z/integraldisplay
dϕe−N tr1
2m2ϕ2ϕi
kϕl
j=1
Z/integraldisplay
dϕe−N1
2m2/summationtext
p,qϕp
qϕq
pϕi
kϕl
j=δi
jδl
k1
Nm2(6)
Setting k=land summing, we find the n=1 term in (5) is equal to (1/z3)δi
j(1/m2).
Just as in any field theory we can associate a Feynman diagram with each of the terms
in (5). For the n=1 term, we have figure VII.4.3. The matrix character of ϕlends itself
naturally to ’t Hooft’s double-line formalism and thus we can speak of quark and gluon
398 | VII. Grand Unification
jl k i
Figure VII.4.3
propagators with a good deal of ease. The Feynman rules are given in figure VII.4.4. We
recognize ϕas the gluon field and (5) as the gluon propagator. Indeed, we can formulate
our problem as follows: Given the bare quark propagator 1 /z, compute the true quark
propagator G(z) with all interaction effects taken into account.
Let us now look at the n=2 term in (5) 1 /z5<ϕi
hϕh
kϕk
lϕl
j>, which we represent in
figure VII.4.5a. With a bit of thought you can see that the index ican be contracted with
k,l,o rj, thus giving rise to figures VII.4.5b, c, d. Summing over color indices, just as in
QCD, we see that the planar diagrams in 5b and 5d dominate the diagram in 5c by a factorN
2. We can take over ’t Hooft’s observation that planar diagrams dominate.
Incidentally, in this example, you see how large Nis essential, allowing us to get rid of
nonplanar diagrams. After all, if I ask you to calculate the density of eigenvalues for sayN=7 you would of course protest saying that the general formula for solving a degree-7
polynomial equation is not even known.
The simple example in figure VII.4.5 already indicates how all possible diagrams could
be constructed. In 5b the same “unit” is repeated, while in 5d the same “unit” is nestedinside a more basic diagram. A more complicated example is shown in 5e. You can convinceyourself that for N=∞ all diagrams contributing to G(z) can be generated by either
“nesting” existing diagrams inside an overarching gluon propagator or “repeating” an
i
kj
l1
z
1
Nm2δi
j δl
k
1
Figure VII.4.4
VII.4. Large NExpansion | 399
i h jl k
(a)
(b)
(c)
(d)
(e)
Figure VII.4.5
existing structure over and over again. T ranslate the preceding sentence into two equations:
“Repeat” (see figure VII.4.6a),
G(z)=1
z+1
z/Sigma1(z)1
z+1
z/Sigma1(z)1
z/Sigma1(z)1
z+...
=1
z−/Sigma1(z)(7)
and “nest” (see figure VII.4.6b),
/Sigma1(z)=1
m2G(z) (8)
400 | VII. Grand Unification
+ =
+
+GΣ
Σ Σ
ΣG(a)
(b)=
Figure VII.4.6
Combining these two equations we obtain a simple quadratic equation for G(z) that we
can immediately solve to obtain
G(z)=m2
2/parenleftBigg
z−/radicalbigg
z2−4
m2/parenrightBigg
(9)
(From the definition of G(z) we see that G(z)→1/zfor large zand thus we choose the
negative root.) We immediately deduce that
ρ(E)=2
πa2/radicalbig
a2−E2 (10)
where a2=4/m2. This is a famous result known as Wigner’s semicircle law.
The Dyson gas
I hope that you are struck by the elegance of the large Nplanar diagram approach. But you
might have also noticed that the gluons do not interact. It is as if we have solved quantum
electrodynamics while we have to solve quantum chromodynamics. What if we have todeal with V( ϕ)=
1
2m2ϕ2+gϕ4? The gϕ4term causes the gluons to interact with each
other, generating horrible diagrams such as the one in figure VII.4.7. Clearly, diagrams
proliferate and as far as I know nobody has ever been able to calculate G(z) using the
Feynman diagram approach.
Happily, G(z) can be evaluated using another method known as the Dyson gas approach.
The key is to write
ϕ=U†/Lambda1U (11)
VII.4. Large NExpansion | 401
Figure VII.4.7
where /Lambda1denotes the NbyNdiagonal matrix with diagonal elements equal to λi,i=
1 ,..., N. Change the integration variable in (3) from ϕtoUand/Lambda1:
Z=/integraldisplay
dU/integraldisplay/parenleftbig
/Pi1idλi/parenrightbig
Je−N/summationtext
kV( λk)(12)
withJthe Jacobian. Since the integrand does not depend on Uwe can throw away the
integral over U. It just gives the volume of the group SU(N) . Does this remind you of
chapter VII.1? Indeed, in (11) Ucorresponds to the unphysical gauge degrees of freedom—
the relevant degrees of freedom are the eigenvalues {λi}. As an exercise you can use the
Faddeev-Popov method to calculate J.
Instead, we will follow the more elegant tack of determining Jby arguing from general
principles. The change of integration variables in (11) is ill defined when any two of the λi’s
are equal, at which point Jmust vanish. (Recall that the change from Cartesian coordinates
to spherical coordinates is ill defined at the north and south poles and indeed the Jacobianin sin θdθdϕ vanishes at θ=0 and π.) Since the λ
i’s are created equal, interchange
symmetry dictates that J=[/Pi1m>n(λm−λn)]β. The power βcan be fixed by dimensional
analysis. With N2matrix elements dϕobviously has dimension λN2while (/Pi1idλi)Jhas
dimension λNλβN(N −1)/2; thus β=2.
Having determined J, let us rewrite (12) as
Z=/integraldisplay
(/Pi1idλi)[/Pi1m>n(λm−λn)]2e−N/summationtext
kV( λk)
=/integraldisplay
(/Pi1idλi)e−N/summationtext
kV( λk)+1
2/summationtext
m/negationslash=nlog(λm−λn)2
(13)
Dyson pointed out that in this form Z=/integraltext
(/Pi1idλi)e−NE(λ 1,...,λN)is just the partition
function of a classical 1-dimensional gas (recall chapter V .2). Think of λi, a real number,
as the position of the ith molecule. The energy of a configuration
E(λ 1,..., λN)=/summationdisplay
kV( λk)−1
2N/summationdisplay
m/negationslash=nlog(λm−λn)2(14)
consists of two terms with obvious physical interpretations. The gas is confined in a
potential well V( x) and the molecules repel4each other with the two-body potential
−(1/N) log(x−y)2. Note that the two terms in Eare of the same order in Nsince each
4Note that this corresponds to the repulsion between energy levels in quantum mechanics.
402 | VII. Grand Unification
sum counts for a power of N. In the large Nlimit (we can think of Nas the inverse
temperature), we evaluate Zby steepest descent and minimize E, obtaining
V/prime(λk)=2
N/summationdisplay
n/negationslash=k1
λk−λn(15)
which in the continuum limit, as the poles in (15) merge into a cut, becomes V/prime(λ)=
2P/integraltext
dμ[ρ(μ)/(λ −μ)], where ρ(μ) is the unknown function we want to solve for and P
denotes principal value.
Defining as before G(z)=/integraltext
dμ[ρ(μ)/(z −μ)] we see that our equation for ρ(μ) can be
written as Re G(λ+iε)=1
2V/prime(λ). In other words, G(z) is a real analytic function with cuts
along the real axis. We are given the real part of G(z) on the cut and are to solve for the
imaginary part. Br ´ezin, Itzykson, Parisi, and Zuber have given an elegant solution of this
problem. Assume for simplicity that V( z ) is an even polynomial and that there is only one
cut (see exercise VII.4.7). Invoke symmetry and, incorporating what we know, postulatethe form
G(z)=1
2/bracketleftBig
V/prime(z)−P(z)/radicalbig
z2−a2/bracketrightBig
withP(z) an unknown even polynomial. Remarkably, the requirement G(z)→1/zfor
largezcompletely determines P(z) . Pedagogically, it is clearest to go to a specific example,
sayV( z )=1
2m2z2+gz4. Since V/prime(z)is a cubic polynomial in z,P(z) has to be a quadratic
(even) polynomial in z. T aking the limit z→∞ and requiring the coefficients of z3and
ofzinG(z) to vanish and the coefficient of 1 /zto be 1 gives us three equations for three
unknowns [namely aand the two unknowns in P(z) ]. The density of eigenvalues is then
determined to be ρ(E)=(1/π)P(E)√
a2−E2.
I think the lesson to take away here is that Feynman diagrams, in spite of their historical
importance in quantum electrodynamics and their usefulness in helping us visualize whatis going on, are vastly overrated. Surely, nobody imagines that QCD, even large NQCD,
will one day be solved by summing Feynman diagrams. What is needed is the analog ofthe Dyson gas approach for large NQCD. Conversely, if a reader of this book manages to
calculate G(z) by summing planar diagrams (after all, the answer is known!), the insight
he or she gains might conceivably be useful in seeing how to deal with planar diagramsin large NQCD.
Field theories in the large Nlimit
A number of field theories have also been solved in the large Nexpansion. I will tell you
about one example, the Gross-Neveu model, partly because it has some of the flavor ofQCD. The model is defined by
S(ψ)=/integraldisplay
d2x⎡
⎣N/summationdisplay
a=1¯ψai/negationslash∂ψa+g2
2N/parenleftBiggN/summationdisplay
a=1¯ψaψa/parenrightBigg2⎤
⎦ (16)
Recall from chapter III.3 that this theory should be renormalizable in (1+1)-dimensional
spacetime. For some finite N, sayN=3, this theory certainly appears no easier to solve
VII.4. Large NExpansion | 403
than any other fully interacting field theory. But as we will see, as N→∞ we can extract
a lot of interesting physics.
Using the identity (A.14) we can rewrite the theory as
S(ψ ,σ)=/integraldisplay
d2x/bracketleftBiggN/summationdisplay
a=1¯ψa(i/negationslash∂−σ)ψa−N
2g2σ2/bracketrightBigg
(17)
By introducing the scalar field σ(x) we have “undone” the four-fermion interaction. (Recall
that we used the same trick in chapter III.5.) You will note that the physics involved issimilar to that behind the introduction of the weak boson to generate the Fermi interaction.Using what we learned in chapters II.5 and IV .3 we can immediately integrate out thefermion fields to obtain an action written purely in terms of the σfield
S(σ)=−/integraldisplay
d2xN
2g2σ2−iN tr log(i/negationslash∂−σ) (18)
Note the factor of Nin front of the tr log term coming from the integration over Nfermion
fields. With the malice of forethought we, or rather Gross and Neveu, have introduced anexplicit factor of 1 /Nin the coupling strength in (16), so that the two terms in (18) both scale
asN. Thus, the path integral Z=/integraltext
Dσe
iS(σ)may be evaluated by the steepest descent or
stationary phase method in the large Nlimit. We simply extremize S(σ) .
Incidentally, we can see the judiciousness of the choice a=1 in large NQCD in the
same way. Integrating out the quarks in (1) we get
S=−/integraldisplay
d4xN
2g2trFμνFμν+Ntr log(i(/negationslash∂−i/negationslashA)−m)
and thus the two terms both scale as Nand can balance each other. The increase in the
number of degrees of freedom has to be offset by a weakening of the coupling.
T o study the ground state behavior of the theory, we restrict our attention to field
configurations σ(x) that do not depend on x. (In other words, we are not expecting
translation symmetry to be spontaneously broken.) We can immediately take over the resultyou got in exercise IV .3.3 and write the effective potential
1
NV( σ)=1
2g(μ)2σ2+1
4πσ2/parenleftbigg
logσ2
μ2−3/parenrightbigg
(19)
We have imposed the condition (1/N)[d2V (σ)/dσ2]|σ=μ=1/g(μ)2as the definition of
the mass scale dependent coupling g(μ) (compare IV .3.18). The statement that V( σ) is
independent of μimmediately gives
1
g(μ)2−1
g(μ/prime)2=1
πlogμ
μ/prime(20)
Asμ→∞ ,g(μ)→0. Remarkably, this theory is asymptotically free, just like QCD. If we
want to, we can work backward to find the flow equation
μd
dμg(μ)=−1
2πg(μ)3+... (21)
The theory in its different incarnations, (16), (17), and (18), enjoys a discrete Z2symme-
try under which ψa→γ5ψaandσ→−σ. As in chapter IV .3, this symmetry is dynamically
404 | VII. Grand Unification
broken by quantum fluctuations. The minimum of V( σ) occurs at σmin=μe1−π/g(μ)2and
so according to (17) the fermions acquire a mass
mF=σmin=μe1−π/g(μ)2(22)
Note that this highly nontrivial result can hardly be seen by staring at (16) and we have no
way of proving it for finite N. In the spirit of the large Napproach, however, we expect that
the fermion mass might be given by mF=μe1−π/g(μ)2+O(1/N2)so that (22) would be a
decent approximation even for say, N=3. Since mFis physically measurable, it better not
depend on μ. You can check that.
This theory also exhibits dimensional transmutation as described in the previous chap-
ter. We start out with a theory with a dimensionless coupling gand end up with a dimen-
sional fermion mass mF. Indeed, any other quantity with dimension of mass would have
to be equal to mFtimes a pure number.
Dynamically generated kinks
I discuss the existence of kinks and solitons in chapter V .6. You clearly understood that the
existence of such objects follows from general considerations of symmetry and topology,rather than from detailed dynamics. Here we have a (1+1)-dimensional theory with a
discrete Z
2symmetry, so we certainly expect a kink, namely a time independent configu-
ration σ(x) (henceforth xwill denote only the spatial coordinate and will no longer label
a generic point in spacetime) such that σ(−∞)=−σmin andσ(+∞)=σmin. [Obviously,
there is also the antikink with σ(−∞)=σmin andσ(+∞)=−σmin.]
At first sight, it would seem almost impossible to determine the precise shape of the
kink. In principle, we have to evaluate tr log[ i/negationslash∂−σ(x) ] for an arbitrary function σ(x) such
thatσ(+∞)=−σ(−∞) (and as I explained in chapter IV .3, this involves finding all the
eigenvalues of the operator i/negationslash∂−σ(x) , summing over the logarithm of the eigenvalues),
and then varying this functional of σ(x) to find the optimal shape of the kink.
Remarkably, the shape can actually be determined thanks to a clever observation.5In
analogy with the steps leading to (IV .3.24) we note that
tr log[i/negationslash∂−σ(x) ]=tr logγ5[i/negationslash∂−σ(x) ]γ5=tr log(−1)[i/negationslash∂+σ(x) ]
and thus up to an irrelevant additive constant
tr log(i/negationslash∂−σ(x)) =1
2tr log[i/negationslash∂−σ(x) ][i/negationslash∂+σ(x) ]
=1
2tr log/braceleftBig
−∂2+iγ1σ/prime(x)−[σ(x) ]2/bracerightBig
(23)
5C. Callan, S. Coleman, D. Gross, and A. Zee, (unpublished). See D. J. Gross, “Applications of the Renormal-
ization Group to High-Energy Physics,” in: R. Balian and J. Zinn-Justin, eds., Methods in Field Theory , p. 247. By
the way, I recommend this book to students of field theory.
VII.4. Large NExpansion | 405
Since γ1has eigenvalues ±i, this is equal to
1
2/braceleftBig
tr log{−∂2+σ/prime(x)−[σ(x) ]2}+tr log {−∂2−σ/prime(x)−[σ(x) ]2}/bracerightBig
but these two terms are equal by parity (space reflection) and hence
tr log[i/negationslash∂−σ(x) ]=tr log{−∂2−σ/prime(x)−[σ(x) ]2}
Referring to (18) we see that S(σ) is the sum of two terms, a term quadratic in σ(x)
and a term that depends only on the combination σ/prime(x)+[σ(x) ]2. But we know that σmin
minimizes S(σ) . Thus, the soliton is given by the solution of the ordinary differential
equation
σ/prime(x)+[σ(x) ]2=σ2
min(24)
namely σ(x)=σmin tanhσminx. The soliton would be observed as an object of size
1/σmin=1/mF. I leave it to you to show that its mass is given by
mS=N
πmF (25)
Precisely as theorized in the last chapter, the ratio mS/mFcomes out to be a pure number,
N/π , as it must.
By an even more clever method that I do not have space to describe, Dashen, Hasslacher,
and Neveu were able to study time dependent configurations of σand determine the mass
spectrum of this model.
Exercises
VII.4.1 Since the number of gluons only differs by one, it is generally argued that it does not make any difference
whether we choose to study the U(N) theory or the SU(N) theory. Discuss how the gluon propagator in
aU(N) theory differs from the gluon propagator in an SU(N) theory and decide which one is easier.
VII.4.2 As a challenge, solve large NQCD in (1+1)-dimensional spacetime. [Hint: The key is that in (1+1)-
dimensional spacetime with a suitable gauge choice we can integrate out the gauge potential Aμ.] For
help, see ’t Hooft, Under the Spell of the Gauge Principle , p. 443.
VII.4.3 Show that if we had chosen to calculate G(z)≡/angbracketleft(1/N) tr(1/z−ϕ)/angbracketright, we would have to connect the two
open ends of the quark propagator. We see that figures VII.4.5b and d lead to the same diagram. Completethe calculation of G(z) in this way.
VII.4.4 Suppose the random matrix ϕis real symmetric rather than hermitean. Show that the Feynman rules
are more complicated. Calculate the density of eigenvalues. [Hint: The double-line propagator can twist.]
VII.4.5 For hermitean random matrices ϕ, calculate
Gc(z,w)≡/angbracketleftbigg1
Ntr1
z−ϕ1
Ntr1
w−ϕ/angbracketrightbigg
−/angbracketleftbigg1
Ntr1
z−ϕ/angbracketrightbigg/angbracketleftbigg1
Ntr1
w−ϕ/angbracketrightbigg
forV( ϕ)=1
2m2ϕ2using Feynman diagrams. [Note that this is a much simpler object to study than the
object we need to study in order to learn about localization (see exercise VI.6.1).] Show that by takingsuitable imaginary parts we can extract the correlation of the density of eigenvalues with itself. For help,see E. Br ´ezin and A. Zee, Phys. Rev. E51: p. 5442, 1995.
406 | VII. Grand Unification
VII.4.6 Use the Faddeev-Popov method to calculate Jin the Dyson gas approach.
VII.4.7 ForV( ϕ)=1
2m2ϕ2+gϕ4, determine ρ(E) .F o rm2sufficiently negative (the double well potential again)
we expect the density of eigenvalues to split into two pieces. This is evident from the Dyson gas picture.Find the critical value m
2
c.F o rm2<m2
cthe assumption of G(z) having only one cut used in the text fails.
Show how to calculate ρ(E) in this regime.
VII.4.8 Calculate the mass of the soliton (25).
VII.5 Grand Unification
Crying out for unification
A gauge theory is specified by a group and the representations the matter fields belong
to. Let us go back to chapter VII.2 and make a catalogue for the SU( 3)⊗SU( 2)⊗U(1)
theory. For example, the left handed up and down quarks are in a doublet/parenleftbiguα
dα/parenrightbig
Lwith
hypercharge1
2Y=1
6. Let us denote this by (3, 2,1
6)L, with the three numbers indicating
how these fields transform under SU( 3)⊗SU( 2)⊗U(1). Similarly, the right handed up
quark is (3, 1,2
3)R. The leptons are (1, 2,−1
2)Land(1, 1,−1)R, where the “1” in the first
entry indicates that these fields do not participate in the strong interaction. Writing it alldown, we see that the quarks and leptons of each family are placed in
(3, 2,1
6)L,(3, 1,2
3)R,(3, 1,−1
3)R,(1, 2,−1
2)L, and (1, 1,−1)R (1)
This motley collection of representations practically cries out for further unification.
Who would have constructed the universe by throwing this strange looking list down?
What we would like to have is a larger gauge group Gcontaining SU( 3)⊗SU( 2)⊗
U(1), such that this laundry list of representations is unified into (ideally) one great big
representation. The gauge bosons in G[but not in SU( 3)⊗SU( 2)⊗U(1)of course] would
couple the representations in (1) to each other.
Before we start searching for G, note that since gauge transformations commute with
the Lorentz group, these desired gauge transformations cannot change left handed fieldsto right handed fields. So let us change all the fields in (1) to left handed fields. Recall fromexercise II.1.9 that charge conjugation changes left handed fields to right handed fieldsand vice versa. Thus, instead of (1) we can write
(3, 2,1
6),(3∗,1 ,−2
3),(3∗,1 ,1
3),(1, 2,−1
2), and (1, 1, 1 ) (2)
We now omit the subscripts LandR: everybody is left handed.
408 | VII. Grand Unification
A perfect fit
The smallest group that contains SU( 3)⊗SU( 2)⊗U(1)isSU( 5). (If you are shaky
about group theory, study appendix B now.) Recall that SU( 5)has 52−1=24 generators.
Explicitly, the generators are represented by 5 by 5 hermitean traceless matrices acting onfive objects we denote by ψ
μwithμ=1, 2, . . . , 5. [These five objects form the fundamental
or defining representation of SU( 5).]
It is now obvious how we can fit SU( 3)andSU( 2)intoSU( 5). Of the 24 matrices that
generate SU( 5), eight have the form/parenleftbigA0
00/parenrightbig
and three the form/parenleftbig00
0B/parenrightbig
, where Arepresents
3 by 3 hermitean traceless matrices (of which there are 32−1=8, the so-called Gell-
Mann matrices) and Brepresents 2 by 2 hermitean traceless matrices (of which there
are 22−1=3, namely the Pauli matrices). Clearly, the former generate an SU( 3)and the
latter an SU( 2). Furthermore, the 5 by 5 hermitean traceless matrix
1
2Y=⎛
⎜⎜⎜⎜⎜⎜⎜⎜⎝−
1
300 0 0
0−1
300 0
00 −1
300
0001
20
000 01
2⎞
⎟⎟⎟⎟⎟⎟⎟⎟⎠(3)
generates a U(1). Without being coy about it, we have already called this matrix the
hypercharge1
2Y.
In other words, if we separate the index μ={α,i}withα=1, 2, 3 and i=4, 5, then
theSU( 3)acts on the index αand the SU( 2)acts on the index i. Thus, the three objects
ψαtransform as a 3-dimensional representation under SU( 3)and hence could be a 3
o ra3∗. Let us choose ψαas transforming as 3; we will see shortly that this is the right
choice with Y/2 given as in (3). The three objects ψαdo not transform under SU( 2)
and hence each of them belongs to the singlet 1 representation. Furthermore, they carryhypercharge −
1
3as we can read off from (3). T o sum up, ψαtransform as (3, 1,−1
3)
under SU( 3)⊗SU( 2)⊗U(1). On the other hand, the two objects ψitransform as 1 under
SU( 3)and 2 under SU( 2), and carry hypercharge1
2; thus they transform as (1, 2,1
2).I n
other words, we embed SU( 3)⊗SU( 2)⊗U(1)intoSU( 5)by specifying how the defining
representation of SU( 5)decomposes into representations of SU( 3)⊗SU( 2)⊗U(1)
5→(3, 1,−1
3)⊕(1, 2,1
2) (4)
T aking the conjugate we see that
5∗→(3∗,1 ,1
3)⊕(1, 2,−1
2) (5)
Inspecting (2), we see that (3∗,1 ,1
3)and(1, 2,−1
2)appear on the list. We are on the right
track! The fields in these two representations fit snugly into 5∗.
This accounts for five of the fields contained in (2); we still have the ten fields
(3, 2,1
6),(3∗,1 ,−2
3), and(1, 1, 1 ) (6)
VII.5. Grand Unification | 409
Consider the next representation of SU( 5)in order of size, namely the antisymmetric
tensor representation ψμν. Its dimension is (5×4)/2=10, precisely the number we want,
if only the quantum numbers under SU( 3)⊗SU( 2)⊗U(1)work out!
Since we know that 5 →(3, 1,−1
3)⊕(1, 2,1
2), we simply (again, see appendix B!) have to
work out the antisymmetric product of (3, 1,−1
3)⊕(1, 2,1
2)with itself, namely the direct
sum of (where ⊗Adenotes the antisymmetric product)
(3, 1,−1
3)⊗A(3, 1,−1
3)=(3∗,1 ,−2
3) (7)
(3, 1,−1
3)⊗A(1, 2,1
2)=(3, 2,−1
3+1
2)=(3, 2,1
6) (8)
and
(1, 2,1
2)⊗A(1, 2,1
2)=(1, 1, 1 ) (9)
[I will walk you through (7): In SU( 3)3⊗A3=3∗(remember εijkfrom appendix B?), in
SU( 2)1⊗A1=1, and in U(1)the hypercharges simply add −1
3−1
3=−2
3.]
Lo and behold, these SU( 3)⊗SU( 2)⊗U(1)representations form exactly the collection
of representations in (6). In other words,
10→(3, 2,1
6)⊕(3∗,1 ,−2
3)⊕(1, 1, 1 ) (10)
The known quark and lepton fields in a given family fit perfectly into the 5∗and 10
representations of SU( 5)!
I have just described the SU( 5)grand unified theory of Georgi and Glashow. In spite of
the fact that the theory has not been directly verified by experiment, it is extremely difficultfor me and for many other physicists not to believe that SU( 5)is at least structurally correct,
in view of the perfect group theoretic fit.
It is often convenient to display the contents of the representation 5
∗and 10, using the
names given to the various fields historically. We write 5∗as a column vector
ψμ=/parenleftBiggψα
ψi/parenrightBigg
=⎛
⎜⎜⎝¯dα
ν
e⎞
⎟⎟⎠(11)
and the 10 as an antisymmetric matrix
ψμν={ψαβ,ψαi,ψij}
=⎛
⎜⎜⎜⎜⎜⎜⎜⎜⎝0¯u−¯udu
−¯u 0¯ud u
¯u−¯u 0du
−d−d−d 0¯e
−u−u−u−¯e 0⎞
⎟⎟⎟⎟⎟⎟⎟⎟⎠(12)
(I suppressed the color indices.)
410 | VII. Grand Unification
Deepening our understanding of physics
Aside from its esthetic appeal, grand unification deepens our understanding of physics
enormously.
1. Ever wondered why electric charge is quantized? Why don’t we see particles with
charge equal to√πtimes the electron’s charge? In quantum electrodynamics, you could
perfectly well write down
L=¯ψ[i(/negationslash∂−i/negationslashA)−m]ψ+¯ψ/prime[i(/negationslash∂−i√π/negationslashA)−m/prime]ψ/prime+... (13)
In contrast, in grand unified theory Aμcouples to a generator of the grand unifying
gauge group, and you know that the generators of any group such as SU(N) (that is
not given by the direct product of U(1)with other groups) are forced by the nontrivial
commutation relations [ Ta,Tb]=ifabcTcto assume quantized values. For example, the
eigenvalues of T3inSU( 2), which depend on the representation of course, must be
multiples of1
2. Within SU( 3)×SU( 2)×U(1), we cannot understand charge quantization:
The generator of U(1)is not quantized. But upon grand unification into SU( 5)[or more
generally any group without U(1)factors] electric charge is quantized.
The result here is deeply connected to Dirac’s remark (chapter IV .4) that electric charge is
quantized if the magnetic monopole exists. We know from chapter V .7 that spontaneouslybroken nonabelian gauge theories such as the SU( 5)theory contain the monopole.
2. Ever wondered why the proton charge is exactly equal and opposite to the electron
charge? This important fact allows us to construct the universe as we know it. Atoms mustbe electrically neutral to some fantastic degree of accuracy for standard cosmology to work;otherwise, electrostatic forces between macroscopic matter would tear the universe apart.
This remarkable fact is nicely incorporated into SU( 5). It is fun to see how it goes.
Evaluating tr Q=0 over the 5
∗implies that 3 Q¯d=−Qe−. I have used the fact that the
strong interaction commutes with electromagnetism and hence quarks with different colorhave the same charge. Now let us calculate the proton charge Q
P:
QP=2Qu+Qd=2(Qd+1)+Qd=3Qd+2=Qe−+2 (14)
IfQe−=− 1, then QP=−Qe−, as is indeed the case!
3. Recall that in electroweak theory we defined tan θ=g1/g2, with the coupling of the
gauge bosons g2Aa
μTa+g1Bμ(Y/2). Since the normalization of Aa
μandBμis fixed by
their respective kinetic energy term, the relative strength of g2andg1is determined by
the normalization of Y/2 relative to T3. Let us evaluate tr T2
3and tr (Y/2)2on the defining
representation 5 : tr T2
3=(1
2)2+(1
2)2=1
2and tr (Y/2)2=(1
3)23+(1
2)22=5
6.
Thus, T3and/radicalbig
3/5(Y/2)are normalized equally. So the correct grand unified combina-
tion is Aa
μTa+Bμ/radicalbig
3/5(Y/2), and therefore tan θ=g1/g2=/radicalbig
3/5o r
sin2θ=3
8(15)
VII.5. Grand Unification | 411
at the grand unification scale. T o compare with the experimental value of sin2θwe would
have to study how the couplings g2andg1flow under the renormalization group down to
low energies. We will postpone this discussion until the next chapter.
Freedom from anomaly
Recall from chapter VII.2 that the key to proving renormalizability of nonabelian gaugetheory is the ability to pass freely between the unitary gauge and the R
ξgauge. The
crucial ingredient is gauge invariance and the resulting Ward-T akahashi identities (seechapter II.7).
Suddenly you start to worry. What about the chiral anomaly? The existence of the
anomaly means that some Ward-T akahashi identities fail to hold. For our theories to makesense, they had better be free from anomalies. I remarked in chapter IV .7 that the historicalname “anomaly” makes it sound like some kind of sickness. Well, in a way, it is.
We should have already checked the SU( 3)⊗SU( 2)⊗U(1)theory for anomalies, but
we didn’t. I will let you do it as an exercise. Here I will show that the SU( 5)theory is healthy.
If theSU( 5)theory is anomaly-free, then a fortiori so is the SU( 3)⊗SU( 2)⊗U(1)theory.
In chapter IV .7 I computed the anomaly in an abelian theory but as I remarked there
clearly all we have to do to generalize to a nonabelian theory is to insert a generator T
a
of the gauge group at each vertex of the triangle diagram in figure IV .7.1. Summing over
the various fermions running around the loop, we see that the anomaly is proportionaltoA
abc(R)≡tr(Ta{Tb,Tc}), where Rdenotes the representation to which the fermions
belong. We have to sum Aabc(R) over all the representations in the theory, remembering
to associate opposite signs to left handed and right handed fermion fields. (It may behelpful to remind yourself of remark 3 in chapter IV .7 and exercise IV .7.6.)
We are now ready to give the SU( 5)theory a health check. First, all fermion fields in (2)
are left handed. Second, convince yourself (simply imagine calculating A
abcfor all possible
abc) that it suffices to set Ta,Tb, andTcall equal to
T≡⎛
⎜⎜⎜⎜⎜⎜⎜⎜⎝2 0 000
0 2 000002 0 0000 −30
000 0 −3⎞
⎟⎟⎟⎟⎟⎟⎟⎟⎠
a multiple of the hypercharge. Let us now evaluate tr T3on the 5∗representation,
trT3|5∗=3(−2)3+2(+3)3=30 (16)
and on the 10,
trT3|10=3(+4)3+6(−1)3+(−6)3=− 30 (17)
An apparent miracle! The anomaly cancels.
412 | VII. Grand Unification
This remarkable cancellation between sums of cubes of a strange list of numbers
suggests strongly, to say the least, that SU( 5)is not the end of the story. Besides, it would
be nice if the 5∗and 10 could be unified into a single representation.
Exercises
VII.5.1 Write down the charge operator Qacting on 5, the defining representation ψμ. Work out the charge
content of the 10 =ψμνand identify the various fields contained therein.
VII.5.2 Show that for any grand unified theory, as long as it is based on a simple group, we have at the unification
scale
sin2θ=/summationtextT2
3/summationtextQ2(18)
where the sum is taken over all fermions.
VII.5.3 Check that the SU( 3)⊗SU( 2)⊗U(1)theory is anomaly-free. [Hint: The calculation is more involved
than in SU( 5)since there are more independent generators. First show that you only have to evaluate
trY{Ta,Tb}and tr Y3, with TaandYthe generators of SU( 2)andU(1), respectively.]
VII.5.4 Construct grand unified theories based on SU( 6),SU( 7),SU( 8), . . . , until you get tired of the game.
People used to get tenure doing this. [Hint: You would have to invent fermions yet to be experimentallydiscovered.]
VII.6 Protons Are Not Forever
Proton decay
Charge conservation guarantees the stability of the electron, but what about the stability of
the proton? Charge conservation allows p→π0+e+. No fundamental principle says that
the proton lives forever, but yet the proton is known for its longevity: It has been aroundessentially since the universe began.
The stability of the proton had to be decreed by an authority figure: Eugene Wigner was
the first to proclaim the law of baryon number conservation. The story goes that whenWigner was asked how he knew that the proton lives forever he quipped, “I can feel it inmy bones.” I take the remark to mean that just from the fact that we do not glow in thedark we can set a fairly good lower bound on the proton’s life span.
As soon as we start grand unifying, we better start worrying. Generically, when we grand
unify we put quarks and leptons into the same representation of some gauge group [see(VII.5.11 and VII.5.12)]. This miscegenation immediately implies that there are gaugebosons transforming quarks into leptons and vice versa. The bag of three quarks knownas the proton could very well get turned into leptons upon the exchange of these gaugebosons. In other words, the proton, the rock on which our world is built upon, may not beforever! Thus, grand unification runs the risk of being immediately falsified.
LetM
Xdenote generically the masses of those gauge bosons transforming quarks into
leptons and vice versa. Then the amplitude for proton decay is of order g2/M2
X, with g
the coupling strength of the grand unifying gauge group, and the proton decay rate /Gamma1is
given by (g2/M2
X)2times a phase space factor controlled essentially by the proton mass
mPsince the pion and positron masses are negligible compared to the proton mass. By
dimensional analysis, we determine that /Gamma1∼(g2/M2
X)2m5P. Since the proton is known to
live for something like at least 1031years, MXhad better be huge compared to the kind of
energy scales we can reach experimentally.
The mass MXis of the same order as the mass scale MGUT at which the grand unified
theory is spontaneously broken down to SU( 3)⊗SU( 2)⊗U(1). Specifically, in the SU( 5)
414 | VII. Grand Unification
μ3
5( )4πg
2
4πg
2
4πg
24πg
2
MGUT1
2
3
Figure VII.6.1
theory, a Higgs field Hμ
νtransforming as the adjoint 24, with its vacuum expectation value
/angbracketleftHμ
ν/angbracketrightequal to the diagonal matrix with elements (−1
3,−1
3,−1
3,1
2,1
2)times some v, can
do the job, as was discussed in chapter IV .6. The gauge bosons in SU( 3)⊗SU( 2)⊗U(1)
remain massless while the other gauge bosons acquire mass MXof order gv.
T o determine MGUT , we apply renormalization group flow to g3,g2, andg1, the cou-
plings of SU( 3),SU( 2), and U(1), respectively. The idea is that as we move up in the
mass or energy scale μthe two asymptotically free couplings g3(μ) andg2(μ) decrease
while g1(μ) increases. Thus, at some mass scale MGUT they will meet and that is where
SU( 3)⊗SU( 2)⊗U(1)is unified into SU( 5)(see figure VII.6.1). Because of the extremely
slow logarithmic running (it should be called walking or even crawling but again for his-torical reasons we are stuck with running) of the coupling constant, we anticipate that theunification mass scale M
GUT will come out to be much larger than any scale we were used
to in particle physics prior to grand unification. In fact, MGUT will turn out to have an
enormous value of the order 1014−15Gev and the idea of grand unification passes its first
hurdle.
Stability of the world implies the weakness of electromagnetism
Using the result of exercise VI.8.1 we obtain (here αS≡g2
3/4πandαGUT≡g2/4πdenote
the strong interaction and grand unification analog of the fine structure constant α,
respectively, with Fthe number of families)
4π
[g3(μ)]2≡1
αS(μ)=1
αGUT+1
6π(4F−33)logMGUT
μ(1)
4π
[g2(μ)]2≡sin2θ(μ)
α(μ)=1
αGUT+1
6π(4F−22)logMGUT
μ(2)
3
54π
[g1(μ)]2≡3
5cos2θ(μ)
α(μ)=1
αGUT+1
6π4FlogMGUT
μ(3)
VII.6. Protons Are Not Forever | 415
Byθ(μ) we mean the value of θat the scale μ.A tμ=MGUT , the three couplings are related
through SU( 5).
We evaluate these equations for some experimentally accessible value of μ, plugging in
measured values of αSandα. With three equations, we not only manage to determine the
unification scale MGUT and coupling αGUT , but we can predict θ. In other words, unless
the ratio g1tog2is precisely right, the three lines in figure VII.6.1 will not meet at one
point.
Note that the number of fermion families Fcontributes equally to (1), (2), and (3). This is
as it should be since the fermions are effectively massless for the purpose of this calculationand do not “know” that the unifying group has been broken into SU( 3)⊗SU( 2)⊗U(1).
These equations are derived assuming that all fermion masses are small compared to μ.
Rearranging these equations somewhat, we find
sin2θ=1
6+5α(μ)
9αS(μ)(4)
sin2θ
α(μ)=1
αS(μ)+1
6π11 logMGUT
μ(5)
1
α(μ)=8
31
αGUT+1
6π/parenleftbigg32
3F−22/parenrightbigg
logMGUT
μ(6)
We obtain in (4) a prediction for sin2θ(μ) independent of MGUT and of the number of
families.
Note that (5) gives the bound
1
α(μ)≥1
6π11 logMGUT
μ(7)
A lower bound on the proton lifetime (and hence on MGUT)translates into an upper bound
on the fine structure constant. Amusingly, the stability of the world implies the weaknessof electromagnetism.
As I noted earlier, plugging in the measured value of α
S, we obtain a huge value for
MGUT . I regard this as a triumph of grand unification: MGUT could have come out to have
a much lower scale, leading to an immediate contradiction with the observed stability of theproton, but it didn’t. Another way of looking at it is that if we are somehow given M
GUT and
αGUT , grand unification fixes the couplings of all three nongravitational interactions! The
point is not that this simplest try at grand unification doesn’t quite agree with experiment:
The miracle is that it works at all.
It is beyond the scope of this book to discuss in detail the comparison of (4), (5), and
(6) with experiment. T o do serious phenomenology, one has to include threshold effects(see exercise VII.6.1), higher order corrections, and so on. T o make a long story short,after grand unified theory came out there was enormous excitement over the possibilityof proton decay. Alas, the experimental lower bound on the proton lifetime was eventuallypushed above the prediction. This certainly does not mean the demise of the notion of
grand unification. Indeed, as I mentioned earlier, the perfect fit is enough to convincemost particle theorists of the essential correctness of the idea. Over the years people haveproposed adding various hypothetical particles to the theory to promote proton longevity.
416 | VII. Grand Unification
The idea is that these particles would affect the renormalization group flow and hence
MGUT . The proton lifetime is actually not the most critical issue. With more accurate
measurements of αSand of θ, it was found that the three couplings do not quite meet at
a point. Indeed, for believers in low energy supersymmetry, part of their faith is foundedon the fact that with supersymmetric particles included, the three coupling constants domeet.
1But skeptics of course can point to the extra freedom to maneuver.
Branching ratios
You may have realized that (1), (2), and (3) are not specific for SU( 5): they hold as long
asSU( 3)⊗SU( 2)⊗U(1)is unified into some simple group (simple so that there is only
one gauge coupling g).
Let us now focus on SU( 5). Recall that we decompose the SU( 5)index μ, which can
take on five values, into two types. In other words, the index μis labeled by {α,i}, where α
takes on three values and itakes on two values. The gauge bosons in SU( 5)correspond to
the 24 independent components of the traceless hermitean field Aμ
ν(μ,ν=1 ,2 ,...,5 )
transforming as the adjoint representation. Focusing on the group theory of SU( 5),w e
will suppress Lorentz indices, spinor indices, etc. Clearly, the eight gauge bosons in SU( 3)
transform an index of type αinto an index of type α, while the three gauge bosons in SU( 2)
transform an index of type iinto an index of type i. Then there is the U(1)gauge boson
that couples to the hypercharge1
2Y. (Of course, you know what I mean by my somewhat
loose language: The SU( 3)gauge bosons transform fields carrying a color index into a
field carrying a color index.)
The fun comes with the gauge bosons Aα
iandAi
α, which transform the index αinto
the index iand vice versa. Since αtakes on three values, and itakes on two values, there
are 6+6=12 such gauge bosons, thus accounting for all the gauge bosons in SU( 5).I n
other words, 24 →(8, 1)+(1, 3)+(1, 1)+(3, 2)+(3∗,2). We will now see explicitly that
the exchange of these bosons between quarks and leptons leads to proton decay.
We merely have to write down the terms in the Lagrangian involving the coupling of
the bosons Aα
iandAi
αto fermions and draw the appropriate Feynman diagrams. I will
go through part of the group theoretic analysis, leaving you to work out the rest. Simplyby contracting indices we see that the boson A
μ
νacting on ψμtakes it to ψνand acting on
ψνρtakes it to ψμρ. Let us look at what A5
αdoes, using your result from exercise VII.5.1.
It takes
ψ5=e−→ψα=¯d (8)
ψαβ=¯u→ψ5β=u (9)
and
ψα4=d→ψ54=e+(10)
1See, e.g., F. Wilczek, in: V . Fitch et al., eds., Critical Problems in Physics , p. 297.
VII.6. Protons Are Not Forever | 417
de+ u u
de+ u u u u
Figure VII.6.2
Thus, the exchange of A5
αgenerates the process (figure VII.6.2) u+d→¯u+e+, leading to
proton decay p(uud) →π0(u¯u)+e+. Observe that while the decay p→π0+e+violates
both baryon number Band lepton number L, it conserves the combination B−L.
In exercise VII.6.2 you will work out the branching ratios for various decay modes. T oo
bad experimentalists have not yet measured them.
Fermion masses
We might hope that with grand unification we would gain new understanding of quarkand lepton masses. Unfortunately, the situation on fermion masses in SU( 5)is muddled,
and to this day nobody understands the origin of quark and lepton masses.
Introducing a Higgs field ϕ
μtransforming as the 5 (as indicated by the notation) we can
write the coupling
ψμCψμνϕν (11)
and
ψμνCψλρϕσεμνλρσ (12)
(withϕνthe conjugate 5∗), reflecting the group theoretic fact (see appendix C) that 5∗⊗10
contains the 5 and 10 ⊗10 contains the 5∗.
Since 5 →(3, 1,−1
3)⊕(1, 2,1
2)we see that this Higgs field is just the natural extension
of the SU( 2)⊗U(1)Higgs doublet (1, 2,1
2). Not wanting to break electromagnetism, we
allow only the electrically neutral fourth component of ϕto acquire a vacuum expectation
value. Setting /angbracketleftϕ4/angbracketright=v , we obtain (up to uninteresting overall constants)
ψαCψα4+ψ5Cψ54/equal1⇒md=me (13)
and
ψαβCψγ5εαβγ/equal1⇒mu/negationslash=0 (14)
418 | VII. Grand Unification
The larger symmetry yields a mass relation md=meat the unification scale; we again
have to apply the renormalization group flow. It is worth noting that the mass relationm
d=mecomes about because as far as the fermions are concerned, SU( 5)has been only
broken down to SU( 4)byϕ. The trouble is that we obtain more or less the same relation
for each of the three families, since most of the running occurs between the unificationscaleM
GUT and the top quark mass so that threshold effects give only a small correction.
Putting in numbers one gets something like
mb
mτ∼ms
mμ∼md
me∼3 (15)
Let us use this to predict the down sector quark masses in terms of the lepton masses.
The formula mb∼3mτworks rather well and provides indirect evidence that there can
only be three families since the renormalization group flow depends on F. The formula
ms∼3mμis more or less in the ballpark, depending on what “experimental” value one
takes for ms. The formula for md, on the other hand, is downright embarrassing. People
mumble something about the first family being so light and hence other effects, such asone-loop corrections might be important. At the cost of making the theory uglier, peoplealso concoct various schemes by introducing more Higgs fields, such as the 45, to givemass to fermions.
Note that in one respect SU( 5)is not as “economical” as SU( 2)⊗U(1), in which the
same Higgs field that gives mass to the gauge bosons also gives mass to the fermions.
The universe is not empty, but almost
I mention in passing another triumph of grand unification: its ability to explain the originof matter in our universe. It has long behooved physicists to understand two fundamentalfacts about the universe: (1) the universe is not empty, and (2) the universe is almost empty.T o physicists, (1) means that the universe is not symmetric between matter and antimatter,that is, the net baryon number N
Bis nonzero; and (2) is quantified by the strikingly small
observed value NB/Nγ∼10−10of the ratio of the number of baryons to the number of
photons.
Suppose we start with a universe with equal quantities of matter and antimatter. For the
universe to evolve into the observed matter dominated universe, three conditions must besatisfied: (1) The laws of the universe must be asymmetric between matter and antimatter.(2) The relevant physical processes had to be out of equilibrium so that there was an arrowof time. (3) Baryon number must be violated.
We know for a fact that conditions (1) and (2) indeed hold in the world: There is
CP violation in the weak interaction and the early universe expanded rapidly. As for
(3), grand unification naturally violates baryon number. Furthermore, while proton decay(suppressed by a factor of 1 /M
2
GUTin amplitude) proceeds at an agonizingly slow rate (for
those involved in the proton decay experiment!), in the early universe, when the Xand
Ybosons are produced in abundance, their fast decays could easily drive baryon number
VII.6. Protons Are Not Forever | 419
violation. The suppression factor 1 /M2
GUTdoes not come in. I have no doubt that eventually
the number 10−10measuring “the amount of dirt in the universe” will be calculated in
some grand unified theory.
Hierarchy
I promised you that the Weisskopf phenomenon would come back to haunt us. That thegrand unification mass scale M
GUT naturally comes out so large counts as a triumph,
but it also leads to a problem known as the hierarchy problem. The hierarchy refers tothe enormous ratio M
GUT/MEW, where MEWdenotes the electroweak unification scale, of
order 102Gev. I will sketch this rather murky subject. Look at the Higgs field ϕrespon-
sible for breaking electroweak theory. We don’t know its renormalized or physical massprecisely, but we do know that it is of order M
EW. Imagine calculating the bare pertur-
bation series in some grand unified theory—the precise theory does not enter into thediscussion—starting with some bare mass μ
0forϕ. The Weisskopf phenomenon tells
us that quantum correction shifts μ2
0by a huge quadratically cutoff dependent amount
δμ2
0∼f2/Lambda12∼f2M2
GUT, where we have substituted for /Lambda1the only natural mass scale
around, namely MGUT , and where fdenotes some dimensionless coupling. T o have the
physical mass squared μ2=μ2
0+δμ2
0come out to be of order M2
EW, something like 28 or-
ders of magnitude smaller than M2
GUT, would require an extremely fine-tuned and highly
unnatural cancellation between μ2
0andδμ2
0. How this could happen “naturally” poses a
severe challenge to theoretical physicists.
Naturalness
The hierarchy problem is closely connected with the notion of naturalness dear to the the-oretical physics community. We naturally expect that dimensionless ratios of parametersin our theories should be of order unity, where the phrase “order unity” is interpreted lib-erally between friends, say anywhere from 10
−2or 10−3to 102or 103. Following ’t Hooft,
we can formulate a technical definition of naturalness: The smallness of a dimensionlessparameter ηwould be considered natural only if a symmetry emerges in the limit η→0.
Thus, fermion masses could be naturally small, since, as you will recall from chapter II.1,a chiral symmetry emerges when a fermion mass is set equal to zero. On the other hand,no particular symmetry emerges when we set either the bare or renormalized mass of ascalar field equal to zero. This represents the essence of the hierarchy problem.
Exercises
VII.6.1 Suppose there are F/primenew families of quarks and leptons with masses of order M/prime. Adopting the crude
approximation described in exercise VI.8.2 of ignoring these families for μbelow M/primeand of treating M/prime
420 | VII. Grand Unification
as negligible for μabove M/prime, run the renormalization group flow and discuss how various predictions,
such as proton lifetime, are changed.
VII.6.2 Work out proton decay in detail. Derive relations between the following decay rates: /Gamma1(p→π0e+),
/Gamma1(p→π+¯ν),/Gamma1(n→π−e+), and/Gamma1(n→π0¯ν).
VII.6.3 Show that SU( 5)conserves the combination B−L. For a challenge, invent a grand unified theory that
violates B−L.
VII.7 SO(10) Unification
Each family into a single representation
At the end of chapter VII.5 we felt we had good reason to think that SU( 5)unification is
not the end of the story. Let us ask if we might be able to fit the 5 and 10∗into a single
representation of a bigger group Gcontaining SU( 5).
It turns out that there is a natural embedding of SU( 5)into the orthogonal SO( 10)
that works,1but to explain that I have to teach you some group theory. The starting point
is perhaps somewhat surprising: We go back to chapter II.3, where we learned that theLorentz group SO( 3, 1), or its Euclidean cousin SO( 4), has spinor representations. We
will now generalize the concept of spinors to d-dimensional Euclidean space. I will work
out the details for deven and leave the odd dimensions as an exercise for you. You might
also want to review appendix B now.
Clifford algebra and spinor representations
Start with an assertion. For any integer nwe claim that we can find 2 nhermitean matrices
γi(i=1, 2, ... ,2n)that satisfy the Clifford algebra
{γi,γj}=2δij (1)
In other words, to prove our claim we have to produce 2 nhermitean matrices γithat
anticommute with each other and square to the identity matrix. We will refer to the γi’s as
theγmatrices for SO( 2n).
Forn=1, it is a breeze: γ1=τ1andγ2=τ2. There you are.
1Howard Georgi told me that he actually found SO( 10)before SU( 5).
422 | VII. Grand Unification
Now iterate. Given the 2 nγ matrices for SO( 2n)we construct the (2 n+2)γmatrices
forSO( 2n+2)as follows
γ(n+1)
j=γ(n)
j⊗τ3=/parenleftBiggγ(n)
j0
0−γ(n)
j/parenrightBigg
,j=1, 2, ... ,2n (2)
γ(n+1)
2n+1=1⊗τ1=/parenleftBigg01
10/parenrightBigg
(3)
γ(n+1)
2n+2=1⊗τ2=/parenleftBigg0−i
i 0/parenrightBigg
(4)
(Throughout this book 1 denotes a unit matrix of the appropriate size.) The superscript in
parentheses is obviously for us to keep track of which set of γmatrices we are talking about.
Verify that if the γ(n)’s satisfy the Clifford algebra, the γ(n+1)’s do as well. For example,
{γ(n+1)
j,γ(n+1)
2n+1}=(γ(n)
j⊗τ3).(1⊗τ1)+(1⊗τ1).(γ(n)
j⊗τ3)
=γ(n)
j⊗{τ3,τ1}=0
This iterative construction yields for SO( 2n)theγmatrices
γ2k−1=1⊗1⊗...⊗1⊗τ1⊗τ3⊗τ3⊗...⊗τ3 (5)
and
γ2k=1⊗1⊗...⊗1⊗τ2⊗τ3⊗τ3⊗...⊗τ3 (6)
with 1 appearing k−1 times and τ3appearing n−ktimes. The γ’s are evidently 2nby 2n
matrices. When and if you feel confused at any point in this discussion you should work
things out explicitly for SO( 4),SO( 6), and so on.
In analogy with the Lorentz group, we define 2 n(2n−1)/2=n(2n−1)hermitean
matrices
σij≡i
2[γi,γj] (7)
Note that σijis equal to iγiγjfori/negationslash=jand vanishes for i=j. The commutation of the σ’s
with each other is thus easy to work out. For example,
[σ12,σ23]=− [γ1γ2,γ2γ3]=−γ1γ2γ2γ3+γ2γ3γ1γ2=− [γ1,γ3]=2iσ13
Roughly speaking, the γ2’s inσ12andσ23knock each other out. Thus, you see that the
1
2σij’s satisfy the same commutation relations as the generators Jij’s ofSO( 2n)(as given
in appendix B). The1
2σij’s represent the Jij’s.
As 2nby 2nmatrices, the σ’s act on an object ψwith 2ncomponents that we will call the
spinor ψ. Consider the unitary transformation ψ→eiωijσijψwithωij=−ωjia set of real
numbers. Then
ψ†γkψ→ψ†e−iωijσijγkeiωijσijψ=ψ†γkψ−iωijψ†[σij,γk]ψ+...
forωijinfinitesimal. Using the Clifford algebra we easily evaluate the commutator as
[σij,γk]=− 2i(δikγj−δjkγi). (Ifkis not equal to either iorjthenγkclearly commutes
VII.7. SO(10) Unification | 423
withσij, and if kis equal to either iorj, then we use γ2
k=1.)We see that the set of objects
vk≡ψ†γkψ,k=1,... ,2ntransforms as a vector in 2 n-dimensional space, with 4 ωijthe
infinitesimal rotation angle in the ijplane:
vk→vk−2(ωkjvj−ωikvi)=vk−4ωkjvj (8)
(in complete analogy to ¯ψγμψtransforming as a vector under the Lorentz group.) This
gives an alternative proof that1
2σijrepresents the generators of SO( 2n).
We define the matrix γFIVE=(−i)nγ1γ2...γ2n, which in the basis we are using has the
explicit form
γFIVE=τ3⊗τ3⊗...⊗τ3 (9)
withτ3appearing ntimes. By analogy with the Lorentz group we define the “left handed”
spinor ψL≡1
2(1−γFIVE)ψand the “right handed” spinor ψR≡1
2(1+γFIVE)ψ, such that
γFIVEψL=−ψLandγFIVEψR=ψR. Under ψ→eiωijσijψ, we have ψL→eiωijσijψLand
ψR→eiωijσijψRsinceγFIVEcommutes with σij. The projection into left and right handed
spinors cut the number of components into halves and thus we arrive at the importantconclusion that the two irreducible spinor representations of SO( 2n)have dimension 2
n−1.
(Convince yourself that the representation cannot be reduced further.) In particular, thespinor representation of SO( 10)is 2
10/2−1=24=16−dimensional. We will see that the
5∗and 10 of SU( 5)can be fit into the 16 of SO( 10).
Embedding unitary groups into orthogonal groups
The unitary group SU( 5)can be naturally embedded into the orthogonal group SO( 10).
In fact, I will now show you that embedding SU(n) intoSO( 2n)is as easy as z=x+iy.
Consider the 2 n-dimensional real vectors x=(x1,... ,xn,y1,... ,yn)and
x/prime=(x/prime
1,... ,x/prime
n,y/prime
1,... ,y/prime
n). By definition, SO( 2n)consists of linear transformations
on these two real vectors leaving their scalar product x/primex=/summationtextn
j=1(x/prime
jxj+y/prime
jyj)invariant.
Now out of these two real vectors we can construct two n-dimensional complex vectors
z=(x1+iy1,... ,xn+iyn)andz/prime=(x/prime
1+iy/prime
1,... ,x/prime
n+iy/prime
n). The group U(n) consists of
transformations on the two n-dimensional complex vectors zandz/primeleaving invariant their
scalar product
(z/prime)∗z=n/summationdisplay
j=1(x/prime
j+iy/prime
j)∗(xj+iyj)
=n/summationdisplay
j=1(x/prime
jxj+y/prime
jyj)+in/summationdisplay
j=1(x/prime
jyj−y/prime
jxj)
In other words, SO( 2n)leaves/summationtextn
j=1(x/prime
jxj+y/prime
jyj)invariant, but U(n) consists of the
subset of those transformations in SO( 2n)that leave invariant not only/summationtextn
j=1(x/prime
jxj+y/prime
jyj)
but also/summationtextn
j=1(x/prime
jyj−y/prime
jxj).
Now that we understand this natural embedding of U(n) intoSO( 2n), we see that the
defining or vector representation of SO( 2n), which we will call simply 2 n, decomposes
424 | VII. Grand Unification
upon restriction to U(n) into the two defining representations of U(n) ,nandn∗; thus
2n→n⊕n∗(10)
In other words, (x1,... ,xn,y1,... ,yn)can be written as (x1+iy1,... ,xn+iyn)and
(x1−iy1,... ,xn−iyn)Note that this is the analog of (VII.5.4) indicating that the defining
representation of SU( 5)decomposes into representations of SU( 3)⊗SU( 2)⊗U(1):
5→(3∗,1 ,1
3)⊕(1, 2,−1
2). (11)
Given the decomposition law (10), we can now figure out how other representations of
SO( 2n)decompose when restricted to the natural subgroup U(n) . The tensor representa-
tions of SO( 2n)are easy, since they are constructed out of the vector representation. [This is
precisely what we did in going from (VII.5.4) to (VII.5.7, 8, and 9).] For example, the adjointrepresentation of SO( 2n), which has dimension 2 n(2n−1)/2=n(2n−1), transforms as
an antisymmetric 2-index tensor 2 n⊗
A2nand so decomposes into
2n⊗A2n→(n⊕n∗)⊗A(n⊕n∗) (12)
according to (10). The antisymmetric product ⊗Aon the right hand side is, of course, to
be evaluated within U(n). For instance, n⊗Anis the n(n−1)/2 representation of U(n) .
In this way, we see that
n(2n−1)→n2−1 (the adjoint)
⊕1 (the singlet)
⊕n(n−1)/2
⊕(n(n−1)/2)∗(13)
As a check, the total dimension of the representations of U(n) on the right hand side adds
up to(n2−1)+1+2n(n−1)/2=n(2n−1). In particular, for SO( 10)⊃SU( 5), we have
45→24⊕1⊕10⊕10∗and of course 24 +1+10+10=45.
Decomposing the spinor
It is more difficult to figure out how the spinor representation of SO( 2n)decompose
upon restriction to U(n) . I give here a heuristic argument that satisfies most physicists,
but certainly not mathematicians. I will just do SO( 10)⊃SU( 5)and let you work out
the general case. The question is how the 16 falls apart. Just from numerology and fromknowing the dimensions of the smaller representations of SU( 5)(1, 5, 10, 15) we see there
are only so many possibilities, some of them rather unlikely, for example, the 16 fallingapart into 16 1’s.
Picture the spinor 16 of SO( 10)breaking up into a bunch of representations of SU( 5).
By definition, the 45 generators of SO( 10)scramble all these representations together. Let
us ask what the various pieces of 45, namely 24 ⊕1⊕10⊕10
∗, do to these representations.
The 24 transform each of the representations of SU( 5)into itself, of course, because they
are the 24 generators of SU( 5)and that is what generators were born to do. The generator
VII.7. SO(10) Unification | 425
1 can only multiply each of these representations by a real number. (In other words, the
corresponding group element multiplies each of the representations by a phase factor.)
What does the 10, which as you recall from chapter VII.5 is represented as an antisym-
metric tensor with two upper indices and hence also known as [2], do to these represen-tations? Suppose the bunch of representations that Sbreaks up into contains the singlet
[0]=1o fSU( 5). The 10 =[2] acting on [0] gives the [2] =10. (Almost too obvious for words!
An antisymmetric tensor of two indices combined with a tensor with no indices is an anti-symmetric tensor of two indices.) What about 10 =[2] acting on [2]? The result is a tensor
with four upper indices. It certainly contains the [4], which is equivalent to [1]
∗=5∗.B u t
look, 1 ⊕10⊕5∗already add up to 16. Thus, we have accounted for everybody. There can’t
be more. So we conclude
S+→[0]⊕[2]⊕[4]=1⊕10⊕5∗(14)
The 5∗and the 10 of SU( 5)fit inside the 16+ofSO( 10)!
We will learn later that the two spinor representations of SO( 10)are conjugate to each
other. Indeed, you may have noticed that I snuck a superscript plus on the letter S. The
conjugate spinor S−breaks up into the conjugate of the representations in (14):
S−→[1]⊕[3]⊕[5]=5⊕10∗⊕1∗(15)
The long lost antineutrino
The fit would be perfect if we introduce one more field transforming as a 1, that is, a singlet
under SU( 5)and hence a fortiori a singlet under SU( 3)⊗SU( 2)⊗U(1). In other words,
this field does not participate in the strong, weak, and electromagnetic interactions, or inplain English, it describes a lepton with no electric charge and is not involved in the knownweak interaction. Thus, this field can be identified as the “long lost” antineutrino field ν
c
L.
This guy does not listen to any of the known gauge bosons.
Recall that we are using a convention in which all fermion fields are left handed, and
hence we have written νc
L. By a conjugate transformation, as explained earlier, this is
equivalent to the right handed neutrino field νR.
SinceνRis anSU( 5)singlet, we can give it a Majorana mass Mwithout breaking SU( 5).
Hence we expect Mto be larger than or of the same order of magnitude as the mass scale
at which SU( 5)is broken, which as we saw in chapter VII.5 is much higher than the mass
scales that have been explored experimentally. This explains why νRhas not been seen.
On the other hand, with the presence of νRwe can have a Dirac mass term m(¯νLνR+
h.c.). Since this term breaks SU( 2)⊗U(1)just like the mass terms for the quarks and
leptons we know, we expect mto be of the same order of magnitude2as the known quark
and lepton masses (which for reasons unknown span an enormous range).
2Explicitly, with νRnow available we can add to the SU( 2)⊗U(1)theory of chapter VII.2 the term f/prime˜ϕ¯νRψL,
where ˜ϕ≡τ2ϕ†. In the absence of any indication to the contrary, we might suppose that f/primeis of the same order
of magntiude as the coupling fthat leads to the electron mass.
426 | VII. Grand Unification
Thus, in the space spanned by (ν,νc)we have the (Majorana) mass matrix
M=/parenleftBigg0m
mM/parenrightBigg
(16)
withM/greatermuchm. Since the trace and determinant of MareMand−m2, respectively, Mhas a
large eigenvalue ∼Mand a small eigenvalue ∼m2/M . A tiny mass ∼m2/M, suppressed
relative to the usual quark and lepton masses by the factor m/M , is naturally generated
for the (observed) left handed neutrino. This rather attractive scenario, known as theseesaw mechanism for obvious reason, was discovered independently by Minkowski andby Glashow, and somewhat later by Y anagida and by Gell-Mann, Ramond, and Slansky.
Again, the tight fit of the 5
∗and the 10 of SU( 5)inside the 16+ofSO( 10)has convinced
many physicists that it is surely right.
A binary code for the world
Given the product form of the γmatrices in (5) and (6), and hence of σij, we can write the
states of the spinor representations as
|ε1ε2...εn/angbracketright (17)
where each of the ε’s takes on the values ±1. For example, for n=1,τ1|+/angbracketright=|−/angbracketright and
τ1|−/angbracketright=|+/angbracketright , while τ2|+/angbracketright=i |−/angbracketright andτ2|−/angbracketright=− i|+/angbracketright . From (9) we see that
γFIVE|ε1ε2...εn/angbracketright=(/Pi1n
j=1εj)|ε1ε2...εn/angbracketright (18)
The right handed spinor S+consists of those states |ε1ε2...εn/angbracketrightwith(/Pi1n
j=1εj)=+ 1, and
the left handed spinor S−those states with (/Pi1n
j=1εj)=− 1. Indeed, the spinor represen-
tations have dimension 2n−1.
Thus, in SO( 10)unification the fundamental quarks and leptons are described by a five-
bit binary code, with states such as |++−−+/angbracketright and|−+−−−/angbracketright . Personally, I find
this a rather pleasing picture of the world.
Let us work out the states explicitly. This also gives me a chance to make sure that
you understand the group theory presented in this chapter. Start with the much simplercase of SO( 4). The spinor S
+consists of |++ /angbracketright and|−− /angbracketright while the spinor S−consists
of|+− /angbracketright and|−+ /angbracketright . As discussed in chapter II.3, SO( 4)contains two distinct SU( 2)
subgroups. Removing a few factors of ifrom the discussion in chapter II.3 we see that
the third generator of SU( 2), call it σ3, can be taken to be either σ12−σ34orσ12+σ34.
The two choices correspond to the two distinct SU( 2)subgroups. We choose (arbitrarily)
σ3=1
2(σ12−σ34). From (5) and (6) we have σ12=iγ1γ2=i(τ 1⊗τ3)(τ2⊗τ3)=−τ3⊗1 and
σ34=− 1⊗τ3, and so σ3=1
2(−τ 3⊗1+1⊗τ3). T o figure out how the four states |++ /angbracketright ,
|−− /angbracketright ,|+− /angbracketright , and|−+ /angbracketright transform under our chosen SU( 2), let us act on them with σ3.
For example,
σ3|++ /angbracketright=1
2(−τ3⊗1+1⊗τ3)|++ /angbracketright=1
2(−1+1)|++ /angbracketright=0
VII.7. SO(10) Unification | 427
and
σ3|−+ /angbracketright=1
2(−τ3⊗1+1⊗τ3)|−+ /angbracketright=1
2(1+1)|−+ /angbracketright=|−+ /angbracketright
Aha, under SU( 2)|++ /angbracketright and|−− /angbracketright are two singlets while |+− /angbracketright and|−+ /angbracketright make up a
doublet.
Note that this is consistent with the generalization of (14) and (15), namely that upon
the restriction of SO( 2n)toU(n) the spinors decompose as
S+→[0]⊕[2]⊕... (19)
and
S−→[1]⊕[3]⊕... (20)
I have not indicated the end of the two sequences: A moment’s reflection indicates that it
depends on whether nis even or odd. In our example, n=2,and thus 2+→[0]⊕[2]=1⊕1
and 2−→[1]=2. Similarly, for n=3, upon the restriction of SO( 6)toU(3),4+→
[0]⊕[2]=1⊕3∗and 4−→[1]⊕[3]=3⊕1. (Our choice of which triplet representation
ofU(3)to call 3 or 3∗is made to conform to common usage, as we will see presently.)
We are now ready to figure out the identity of each of the 16 states such as |++−−+/angbracketright
inSO( 10)unification. First of all, (18) tells us that under the subgroup SO( 4)⊗SO( 6)of
SO( 10)the spinor 16+decomposes as (since /Pi15
j=1εj=+ 1 implies ε1ε2=ε3ε4ε5)
16+→(2+,4+)⊕(2−,4−) (21)
We identify the natural SU( 2)subgroup of SO( 4)as the SU( 2)of the electroweak interac-
tion and the natural SU( 3)subgroup of SO( 6)as the color SU( 3)of the strong interaction.
Thus, according to the preceding discussion, (2+,4+)are the SU( 2)singlets of the stan-
dardU(1)⊗SU( 2)⊗SU( 3)model, while (2−,4−)are the SU( 2)doublets. Here is the
lineup (all fields being left handed as usual):
SU( 2)doublets:
ν=|−+−−−/angbracketright
e−=|+−−−−/angbracketright
u=|−+++−/angbracketright ,|−++−+/angbracketright , and |−+−++/angbracketright
d=|+−++−/angbracketright ,|+−+−+/angbracketright , and |+−−++/angbracketright
SU( 2)singlets:
νc=|+++++/angbracketright
e+=|−−+++/angbracketright
uc=|+++−−/angbracketright ,|++−+−/angbracketright , and |++−−+/angbracketright
dc=|−−+−−/angbracketright ,|−−−+−/angbracketright , and |−−−−+/angbracketright .
I assure you that this is a lot of fun to work out and I urge you to reconstruct this
table without looking at it. Here are a few hints if you need help. From our discussion
428 | VII. Grand Unification
ofSU( 2)I know that ν=|−+ ε3ε4ε5/angbracketrightande−=|+− ε3ε4ε5/angbracketright, but how do I know that
ε3=ε4=ε5=− 1? First, I know that ε3ε4ε5=− 1. I also know that 4−→3⊕1 upon
restricting SO( 6)to color SU( 3). Well, of the four states |−−−/angbracketright ,|++−/angbracketright ,|+−+/angbracketright ,
and|−++/angbracketright the “odd man out” is clearly |−−−/angbracketright . By the same heuristic argument,
among the 16 possible states |+++++/angbracketright is the “odd man out” and so must be νc.
There are lots of consistency checks. For example, once I identify ν=|−+−−−/angbracketright ,
e−=|+−−−−/angbracketright , and νc=|+++++/angbracketright , I can figure out the electric charge Q,
which, since it transforms as a singlet under color SU( 3), must have the value Q=
aε1+bε2+c(ε 3+ε4+ε5)when acting on the state |ε1ε2ε3ε4ε5/angbracketright. The constants a,b, and
ccan be determined from the three equations Q(ν)=−a+b−3c=0,Q(e−)=− 1, and
Q(νc)=0. Thus, Q=−1
2ε1+1
6(ε3+ε4+ε5).
Living in the computer age, I find it intriguing that the fundamental constituents
of matter are coded by five bits. You can tell your condensed matter colleagues thattheir beloved electron is composed of the binary strings +−−−− and−−+++ .A n
intriguing possibility
3suggests itself, that quarks and leptons may be composed of five
different species of fundamental fermionic objects. We construct composites, writing a+if that species is present, and a −if it is absent. For example, from the expression for
Qgiven above, we see that species 1 carries electric charge −
1
2, species 2 is neutral, and
species 3, 4, and 5 carry charge1
6. A more or less concrete model can even be imagined by
binding these fundamental fermionic objects to a magnetic monopole.
I emphasize that particles transforming in 16−, such as |+−+++/angbracketright , have not been
observed experimentally.
A speculation on the origin of families
One of the great unsolved puzzles in particle physics is the family problem. Why do quarksand leptons come in three generations {ν
e,e,u,d},{νμ,μ,c,s}, and{ντ,τ,t,b}? The way
we incorporate this experimental fact into our present day theory can only be describedas pathetic: We repeat the fermionic sector of the Lagrangian three times without anyunderstanding whatsoever. Three generations living together gives rise to a nagging familyproblem.
Our binary code view of the world suggests a wildly speculative (perhaps too speculative
to mention in a textbook?) approach to the family problem: We add more bits. T o me,a reasonable possibility is to “hyperunify” into an SO( 18)theory, putting all fermions
into a single spinorial representation S
+=256+, which upon the breaking of SO( 18)to
SO( 10)⊗SO( 8)decomposes as
256+→(16+,8+)⊕(16−,8−) (22)
We have a lot of 16+’s. Unhappily, we see that group theory [see also (21)] dictates that we
also get a bunch of unwanted 16−’s. One suggestion is that Nature might repeat the trick
3For further details, see F. Wilczek and A. Zee, Phys. Rev. D25: 553, 1982, Section IV .
VII.7. SO(10) Unification | 429
She uses with color SU( 3), whose strong force confines fields that are not color singlets
(chapter VII.3). Interestingly, we can exploit a striking feature of SO( 8), which some people
regard as the most beautiful of all groups. In particular, the two spinorial representations8
±have the same dimension as the vectorial representation 8v(the equation 2n−1=2n
has the unique solution n=4). There is a transformation that cyclically rotates these three
representations 8+,8−, and 8vinto each other (in the jargon, the group SO( 8)admits an
outer automorphism). Thus, there exists a subgroup SO( 5)ofSO( 8)such that when we
break SO( 8)into that SO( 5)8+behaves like 8vwhile 8−behaves like a spinor, namely
8+→5⊕1⊕1⊕1 and 8−→4⊕4∗(23)
If we call this4SO( 5)hypercolor and assumes that the strong force associated with it con-
fines all fields that are not hypercolor singlets, then only three 16+’s remain! Unfortunately,
as the relevant physics occurs in the energy regime above grand unification, our knowl-edge of the dynamics of symmetry breaking is far too paltry for us to make any furtherstatements.
Charge conjugation
The product ⊗notation we use here allows us to construct the conjugation matrix Cexplic-
itly. By definition C−1σ∗
ijC=−σij(so that Cchanges eiθijσijinto its complex conjugate.)
From (2), (3), and (4) we see that we can construct
C(n+1)={C(n)⊗τ1ifnodd
C(n)⊗τ2 ifneven(24)
You can check that this gives C−1γ∗
jC=(−1)nγjand hence the desired result.
Explicitly, Cis a direct product of an alternating sequence of τ1andτ2and so we deduce
an important property. Acting on |ε1ε2...εn/angbracketright,Cflips the sign of all the ε’s. Thus C
changes the sign of (/Pi1n
j=1εj)fornodd, and does not for neven. For nodd, the two
spinor representations S+andS−are conjugates of each other, while for neven, they
are conjugates of themselves, or in other words, they are real. This can also be seendirectly from C
−1γFIVEC=(−1)nγFIVE. You can check this with all the explicit examples
we have encountered: SO( 2),SO( 4),SO( 6),SO( 8),SO( 10), and SO( 18). See also
exercise VII.7.3.
Anomalies
What about anomalies in SO( 2n)grand unification? According to the discussion in chap-
ter VII.5 we have to evaluate Aijklmn≡tr(Jij{Jkl,Jmn})over the fermion representation.
4The reader savvy with group theory would recognize that SO( 5)is isomorphic with the symplectic group
Sp( 4)and that the Dynkin diagram of SO( 8)is the most symmetric of all.
430 | VII. Grand Unification
Applying an SO( 2n)transformation Jij→OTJijOwe see easily that Aijklmnis an invari-
ant tensor. Can we construct an invariant 6-index tensor with the appropriate symmetryproperties (e.g., A
ijklmn=−Ajiklmn)inSO( 2n)? We can’t, except in SO( 6), for which we
haveεijklmn. Thus, Aijklmnvanishes except in SO( 6), where it is proportional to εijklmn.
An elegant one line proof that any grand unified theory based on SO( 2n)forn/negationslash=3 is free
from anomaly!
The cancellation of the anomaly between 5∗and 10 at the end of chapter VII.5 doesn’t
seem so miraculous any more. Miracles tend to fade away as we gain deeper understanding.
Amusingly, by discussing a physics question, namely whether a gauge theory is renor-
malizable or not, we have discovered a mathematical fact. What is so special about SO( 6)?
See exercise VII.7.5.
Exercises
VII.7.1 Work out the Clifford algebra in d-dimensional space for dodd.
VII.7.2 Work out the Clifford algebra in d-dimensional Minkowski space.
VII.7.3 Show that the Clifford algebra for d=4kand for d=4k+2 have somewhat different properties. (If you
need help with this and the two preceding exercises, look up F. Wilczek and A. Zee, Phys. Rev. D25: 553,
1982.)
VII.7.4 Discuss the Higgs sector of the SO( 10). What do you need to give mass to the quarks and leptons?
VII.7.5 The group SO( 6)has 6(6−1)/2=15 generators. Notice that the group SU( 4)also has 42−1=15
generators. Substantiate your suspicion that SO( 6)andSU( 4)are isomorphic. Identify some low
dimensional representations.
VII.7.6 Show that (unfortunately) the number of families we get in SO( 18)depends on which subgroup of SO( 8)
we take to be hypercolor.
VII.7.7 If you want to grow up to be a string theorist, you need to be familiar with the Dirac equation in various
dimensions but especially in 10. As a warm up, study the Dirac equation in 2-dimensional spacetime.Then proceed to study the Dirac equation in 10-dimensional spacetime.
Part VIII Gravity and Beyond
This page intentionally left blank
VIII.1Gravity as a Field Theory
and the Kaluza-Klein Picture
Including gravity
Field theory texts written a generation ago typically do not even mention gravity. The
gravitational interaction, being so much weaker than the other three interactions, wassimply not included in the education of particle physicists. The situation has changed witha vengeance: The main drive of theoretical high energy physics today is the unification ofgravity with the other three interactions, with string theory the main candidate for a unifiedtheory.
From a course on general relativity you would have learned about the Einstein-Hilbert
action for gravity
S=1
16πG/integraldisplay
d4x√−gR≡/integraldisplay
d4x√−gM2
PR (1)
where g=detgμνdenotes the determinant of the curved metric gμνof spacetime, Ris
the scalar curvature, and Gis Newton’s constant. Let me remind you that the Riemann
curvature tensor
Rλ
μνκ=∂ν/Gamma1λ
μκ−∂κ/Gamma1λ
μν+/Gamma1σ
μκ/Gamma1λ
νσ−/Gamma1σ
μν/Gamma1λ
κσ(2)
is constructed out of the Riemann-Christoffel symbol (recall chapter I.11):
/Gamma1λ
μν=1
2gλρ(∂νgρμ+∂μgρν−∂ρgμν) (3)
The Ricci tensor is defined by Rμκ=Rν
μνκand the scalar curvature by R=gμνRμν. Varying
Sgives us1the Einstein field equation
Rμν−1
2gμνR=− 8πGTμν (4)
The Einstein-Hilbert action is uniquely determined if we require the action to be coor-
dinate invariant and to involve two powers of spacetime derivative. As you can see from
1See, e.g., S. Weinberg, Gravitation and Cosmology , p. 364.
434 | VIII. Gravity and Beyond
(2) and (3) the scalar curvature Rinvolves two powers of derivative and the dimensionless
fieldgμνand thus has mass dimension 2. Hence G−1must have mass dimension 2. The
second form in (1) emphasizes this point and is often preferred in modern work on grav-ity. (The modified Planck mass M
P≡1/√
16πG differs from the usual Planck mass by a
trivial factor, much like the relation between hand/planckover2pi.)
The theory sprang from Einstein’s profound intuition regarding the curvature of space-
time and is manifestly formulated in terms of geometric concepts. In many textbooks,Einstein’s theory is developed, and rightly so, in purely geometric terms.
On the other hand, as I hinted back in chapter I.6, gravity can be treated on the same
footing as the other interactions. After all, the graviton may be regarded as just anotherelementary particle like the photon. The action (1), however, does not look anything likethe field theories we have studied thus far. I will now show you that in fact it does have thesame kind of structure.
Gravity as a field theory
Let us write gμν=ημν+hμν, where ημνdenotes the flat Minkowski metric and hμνthe
deviation from the flat metric. Expand the action in powers of hμν. In order not to drown
in a sea of Lorentz indices, let us suppress them for a first go-around. Merely from thefact that the scalar curvature Rinvolves two derivatives ∂in its definition, we see that the
expansion must have the schematic form
S=/integraldisplay
d4x1
16πG(∂h∂h +h∂h∂h +h2∂h∂h+...) (5)
after dropping total divergences. As I remarked in chapter I.11, the field hμν(x) describes
a graviton in flat space and is to be treated like any other field. The first term ∂h∂h , which
governs how the graviton propagates, is conceptually no different than the first term in theaction for a scalar field ∂ϕ∂ϕ or for the photon field ∂A∂A . The terms cubic and higher in
hdetermine the interaction of the graviton with itself.
The Einstein-Hilbert action in the weak field expansion is structurally reminiscent of the
Y ang-Mills action, which may be written in schematic form as S=/integraltext
d
4x(1/g2)(∂A∂A +
A2∂A+A4). As I explained in chapter IV .5, we understand the self interaction of the Y ang-
Mills bosons physically: The bosons themselves carry the charge to which they couple.
We can understand the self interaction of the graviton similarly: The graviton couples toanything carrying energy and momentum, and it certainly carries energy and momentum.In contrast, the photon does not couple to itself.
We say that Y ang-Mills and Einstein theories are nonlinear, while Maxwell theory is
linear. The former are hard, the latter easy.
But while the Y ang-Mills action terminates, the Einstein-Hilbert action, because of the
presence of√
−g and of the inverse of gμν, is an infinite series in the graviton field hμν.
The other major difference is that while Y ang-Mills theory is renormalizable, gravity is
notoriously nonrenormalizable, as we argued by dimensional analysis in chapter III.2. Weare now in a position to see this explicitly. Consider the self energy correction to the graviton
VIII.1. Gravity as a Field Theory | 435
(a) (b)
Figure VIII.1.1
propagator shown in figure VIII.1.1a. We see from the second term in (5) that the three-
graviton coupling involves two powers of momentum. Thus the Feynman integral goes as/integraltext
d4k(kkkk/k2k2), with four powers of kin the numerator from the two vertices and four
powers in the denominator from the propagators. T aking out two powers of momentum toextract the coefficient of ∂h∂h , we see that the correction to 1 /G is quadratically divergent.
Because of the explicit powers of momentum in the coupling, the divergence gets worseand worse as we go to higher and higher order. Compare figure VIII.1.1b to 1a: We havethree more propagators, worth ∼1/k
6, and one more loop integration/integraltext
d4k, but two more
vertices ∼k4. The degree of divergence goes up by 2. Of course we already knew all this
by dimensional analysis.
As mentioned in chapter I.11 the fundamental definition
Tμν(x)=−2√−gδSM
δgμν(x)
tells us that coupling of the graviton to matter (in the weak field limit) can be included by
adding the term
−/integraldisplay
d4x1
2hμνTμν(6)
to the action, where Tμνstands for the (flat spacetime) stress-energy tensor of all the matter
fields of the world, a matter field being any field that is not the graviton field. Thus, withthe inclusion of matter (5) is modified schematically
2to
S=/integraldisplay
d4x[1
16πG(∂h∂h +h∂h∂h +h2∂h∂h+...)+(hT+...)] (7)
In chapter IV .5 I noted that we can bring Y ang-Mills theory into the same convention
commonly used in Maxwell theory by a trivial rescaling A→gA. Similarly, we can also
bring Einstein theory into the same convention by rescaling the graviton field hμν→√
Ghμνso that the action becomes (to ease writing we absorb 16 πintoGwhenever we
feel like it)
S=/integraldisplay
d4x (∂h∂h +√
Gh∂h∂h +Gh2∂h∂h+...+√
GhT )
2If this is to represent an expansion of Sin powers of h, then strictly speaking, if we display the terms cubic
and quartic in the Einstein-Hilbert action, we should also display the contribution coming from the terms of
higher order in hcontained in Tμν(x)=−(2/√−g)δSM/δgμν(x).
436 | VIII. Gravity and Beyond
We see explicitly that√
16πG=1/MPmeasures the strength of the graviton coupling to
itself and to all other fields. Once again, the enormity of MP(compared to the scale of the
strong interaction, say) indicates the feebleness of gravity.
Here we expanded gμνaround a flat metric but we could just as well expand gμν=
¯gμν+hμν, with ¯gμνa curved metric, that of a black hole (see chapter V .7) for instance.
Determining the weak field action
After this index free survey we are ready to tackle the indices. We would like to determine
the first term ∂h∂h in (7) so that we can obtain the graviton propagator. Thus, we have to
expand the action S≡M2
P/integraltext
d4x√−ggμνRμνup to and including order h2. From (2) and
(3) we see that the Ricci tensor Rμνstarts in O(h) so that it suffices to evaluate√−ggμνto
O(h) . That’s easy: As we have already seen in chapter I.11, g=− [1+ημνhμν+O(h2)] and
gμν=ημν−hμν+O(h2)so that√−ggμν=ημν−hμν+1
2ημνh+O(h2), where we have
defined h≡ημνhμν. We now must calculate RμνtoO(h2), a straightforward but tedious
task starting from (2) and (3).
In line with the spirit of this book, which is to avoid tedious calculation whenever
possible, I will now show you how to get around this. We invoke symmetry considerations!
Under a general coordinate transformation xμ→x/primeμ=xμ−εμ(x) the metric changes
tog/primeμν=(∂x/primeμ/∂xσ)(∂x/primeν/∂xτ)gστ. Plugging in gμν=ημν−hμν+..., lowering the in-
dices (with ημνto this order), and using (∂x/primeμ/∂xσ)=δμ
σ−∂σεμ, we find, treating ∂μενas
of the same order as hμν:
h/prime
μν=hμν+∂μεν+∂νεμ (8)
Note the structural similarity to the electromagnetic gauge transformation A/prime
μ=Aμ−
∂μ/Lambda1. Very nice! We will explore the sense in which gravity can be regarded as a gauge
theory in more detail later.
We are looking for the terms in the action quadratic in hand quadratic in ∂. Lorentz
invariance tells us that there are four possible terms (T o see this, first write down termswith the indices on the two ∂matching, then the terms with the index on a ∂matching an
index on an h, and so on):
S=/integraldisplay
d4x(a∂ λhμν∂λhμν+b∂λhμ
μ∂λhνν+c∂λhλν∂μhμν+dhλ
λ∂μ∂νhμν)
with four unknown constants a,b,c, andd. Now vary Swithδhμν=∂μεν+∂νεμ, inte-
grating by parts freely. For example,
δ(∂λhμν∂λhμν)=2[∂λ(2∂μεν)](∂λhμν)“=”4εν∂2∂μhμν
Since there are three objects linear in h, linear in ε, and cubic in ∂(namely εν∂2∂νhand
εν∂ν∂λ∂μhλμin addition to the one already shown) the condition δS=0 gives three equa-
tions, just enough to fix the action up to an overall constant, corresponding to Newton’sconstant. The invariant combination turns out to be
I≡1
2∂λhμν∂λhμν−1
2∂λhμ
μ∂λhνν−∂λhλν∂μhμν+∂νhλ
λ∂μhμν (9)
VIII.1. Gravity as a Field Theory | 437
Thus, even if we had never heard of the Einstein-Hilbert action we could still determine
the action for gravity in the weak field limit by requiring that the action be invariant underthe transformation (8). This is hardly surprising since coordinate invariance determinesthe Einstein-Hilbert action. Still, it is nice to construct gravity “from scratch.”
Referring to (6), we can now write the weak field expansion of Sas
Swfg=/integraldisplay
d4x/parenleftbigg1
32πGI−1
2hμνTμν/parenrightbigg
without having to expand RtoO(h2). The coefficient of Iis fixed by the requirement that
we reproduce the usual Newtonian gravity (see later).
The graviton propagator
As we anticipated in (5) the action Swfg indeed has the same quadratic structure of all
the field theories we have studied, and so as usual the graviton propagator is just theinverse of a differential operator. But just as in Maxwell and Y ang-Mills theories the relevantdifferential operator in Einstein-Hilbert theory does not have an inverse because of the“gauge invariance” in (8).
No problem. We have already developed the Faddeev-Popov method to deal with this
difficulty. In fact, for my limited purposes here, to derive the graviton propagator in flatspacetime, I don’t even need the full-blown Faddeev-Popov formalism with ghosts and all.
3
Indeed, recall from chapter III.4 that for the Feynman gauge (ξ =1)we simply add (∂A)2
to the invariant
1
2FμνFμν=∂μAν(∂μAν−∂νAμ)“=”−Aμημν∂2Aν−(∂A)2
thus canceling the last term. Inverting the differential operator −ημν∂2we obtain the
photon propagator in the Feynman gauge −iημν/k2. We play the same “trick” for gravity.
After staring at
I=1
2∂λhμν∂λhμν−1
2∂λhμ
μ∂λhνν−∂λhλν∂μhμν+∂νhλ
λ∂μhμν
for a while, we see that by adding (∂μhμν−1
2∂νhλ
λ)2we can knock off the last two terms in
Iso that Swfgeffectively becomes
Swfg=/integraldisplay
d4x1
2/bracketleftbigg1
32πG/parenleftbigg
∂λhμν∂λhμν−1
2∂λh∂λh/parenrightbigg
−hμνTμν/bracketrightbigg
(10)
In other words, the freedom in choosing hμνin (8) allows us to impose the so-called
harmonic gauge condition
∂μhμ
ν=1
2∂νhλλ (11)
(the linearized version of ∂μ(√−ggμν)=0.)
3This is because (8) does not involve the field hμν, just as in the Maxwell case but unlike the Y ang-Mills
case. Since we do not intend to calculate loop diagrams in quantum gravity, we do not need the full power of the
Faddeev-Popov method.
438 | VIII. Gravity and Beyond
Writing (10) in the form
S=1
32πG/integraldisplay
d4x/bracketleftBig
hμνKμν;λσ(−∂2)hλσ+O(h3)/bracketrightBig
we see that we have to invert the matrix
Kμν;λσ≡1
2(ημληνσ+ημσηνλ−ημνηλσ)
regarding μν andλσ as the two indices. Note that we have to maintain the symmetry of
hμν. In other words, we are dealing with matrices acting in a linear space spanned by
symmetric two-index tensors. Thus, the identity matrix is actually
Iμν;λσ≡1
2(ημληνσ+ημσηνλ)
You can check that Kμν;λσKλσ
;ρω=Iμν;ρωso that K−1=K. Thus, in the harmonic gauge
the graviton propagator in flat spacetime is given by (scaling out Newton’s constant)
Dμν,λσ(k)=1
2ημληνσ+ημσηνλ−ημνηλσ
k2+iε(12)
Newton from Einstein
Varying (10) with respect to hμνwe obtain the Euler-Lagrange equation of motion4
1
32πG(−2∂2hμν+ημν∂2h)−Tμν=0. T aking the trace, we find ∂2h=16πGT (withT≡
ημνTμν)and so we obtain5
∂2hμν=− 16πG(Tμν−1
2ημνT) (13)
In the static limit, T00is the dominant component6of the stress-energy tensor and (13)
reduces to /vector∇2φ=4πGT00upon recalling from chapter I.5 that the Newtonian gravitational
potential φ≡1
2h00. We have just derived Poisson’s equation for φ.
Incidentally, this suggests another way of avoiding the tedious task of expanding the
Einstein-Hilbert action (and hence R)toO(h2)if you are willing to accept the Einstein
field equation (4) as given. You need expand Rμνonly to O(h) to obtain (13) from (4), and
from (13) you can reconstruct the action to O(h2). Indeed, from (2) and (3) you easily get
Rμν=1
2(−∂2hμν+∂μ∂λhλ
ν+∂ν∂λhλ
μ−∂μ∂νhλ
λ)+O(h2)→−1
2∂2hμν+O(h2)
with the further simplification in harmonic gauge. But this is not quite fair since con-
siderable technology7(Palatini identity and all the rest) is needed to derive (4) from (1).
4Note that the flat spacetime energy momentum conservation ∂μTμν=0 together with the equation of motion
implies ∂2(∂μhμν−1
2∂νh)=0.
5Thus, the Einstein equation in vacuum Rμν=0 reduces to ∂2hμν=0; hence the name “harmonic.”
6Note that, in contrast to T00,h00does not dominate the other components of hμν.
7See S. Weinberg, Gravitation and Cosmology , pp. 290 and 364.
VIII.1. Gravity as a Field Theory | 439
Einstein’s theory and the deflection of light
Consider two particles with stress-energy tensors Tμν
(1)andTμν
(2)respectively interacting via
the exchange of a graviton. The scattering amplitude is then (up to some overall constantnot essential for our purposes here) given by
GTμν
(1)Dμν,λσ(k)Tλσ
(2)=G
2k2(2Tμν
(1)T(2)μν−T(1)T(2))
For nonrelativistic matter T00is much larger than the other components T0jandTij(as
I have just remarked), so the scattering amplitude between two lumps of nonrelativisticmatter (say, the earth and you) is proportional to
G
2k2(2T00
(1)T00
(2)−T00
(1)T00
(2))=G
2k2T00
(1)T00
(2)
As explained way back in chapters I.4 and I.5, the interaction potential is given by the
Fourier transform of the scattering amplitude, namely
G/integraldisplay/integraldisplay
d3xd3x/primeT(1)00(x)T(2)00(x/prime)/integraldisplay
d3kei/vectork.(/vectorx−/vectorx/prime)1
/vectork2
and thus for two well-separated objects we recover the Newtonian potential GM(1)M(2)/r.
We are now able to address the issue raised at the end of chapter I.5. Suppose a particle
theorist, Dr. Gravity, wants to propose a theory of gravity to rival Einstein’s theory. Dr. Gclaims that gravity is due to the exchange of a spin 2 particle with a teeny mass m
Gcoupled
to the stress-energy tensor Tμν. In chapter I.5 we worked out the propagator of a massive
spin 2 particle, namely
Dspin 2
μν,λσ(k)=1
2(GμλGνσ+GμσGνλ−2
3GμνGλσ)/(k2−m2
G+iε)
withGμν=ημν−kμkν/m2
G(after a trivial notational adjustment). Since the particle is
coupled to a conserved source kμTμν=0 we can replace Gμνbyημν. Thus, in the limit
mG→0 we have the propagator
Dspin 2
μν,λσ(k)=1
2ημληνσ+ημσηνλ−2
3ημνηλσ
k2+iε(14)
Compare this with (12). Dr. G’s propagator differs from Einstein’s:2
3versus 1. Remarkably,
gravity is not generated by an almost massless spin 2 particle. The “2
3discontinuity”
between (12) and (14) was discovered in 1970 independently by Iwasaki, by van Dam and
Veltman, and by Zakharov.
In Dr. G’s theory (with his own gravitational coupling GG), the interaction between two
particles is given by
GGTμν
(1)Dμν,λσ(k)Tλσ
(2)=GG
2k2(2Tμν
(1)T(2)μν−2
3T(1)T(2))
For two lumps of nonrelativistic matter this becomes
GG
2k2(2T00
(1)T00
(2)−2
3T00
(1)T00
(2))=4
3GG
2k2T00
(1)T00
(2)
Dr. G simply takes his GG=3
4Gand his theory passes all experimental tests.
440 | VIII. Gravity and Beyond
But wait! There is also the famous 1919 observation of the deflection of starlight by
the sun, and the photon is definitely not a lump of nonrelativistic matter. Indeed, recallfrom chapter I.11 (or from your course on electromagnetism) that T≡T
μ
μvanishes for
the photon. Thus, taking Tμν
(1)andTμν
(2)to be the stress-energy tensor of the sun and of the
photon respectively, Einstein would have for the scattering amplitude (G/2k2)2Tμν
(1)T(2)μν
while Dr. G would have (GG/2k2)2Tμν
(1)T(2)μν=3
4(G/2k2)2Tμν
(1)T(2)μν. Dr. G would have
predicted a deflection angle of 3 GM/R instead of 4 GM/R (with MandRthe mass
and radius of the sun). On the Brazilian island of Sobral in 1919 Einstein triumphedover Dr. G.
As explained in chapter I.5, while a massive spin 2 particle has 5 degrees of freedom the
massless graviton has only 2. (I give an analysis of the helicity ±2 structure of one graviton
exchange in appendix 2.) The 5 degrees of freedom may be thought of as consisting of thehelicity ±2 degrees of freedom we want plus 2 helicity ±1 and a helicity 0 degrees of
freedom. The coupling of the helicity ±1 degrees of freedom vanishes because k
μTμν=0.
Thus, effectively, we are left with an extra scalar coupling to the trace T≡ημνTμνof the
stress-energy tensor; as we can see plainly the discrepancy indeed resides in the last termof (12) and (14).
You should be disturbed that a measurement of the deflection of starlight can show
that a physical quantity, the graviton mass m
G, is mathematically zero rather than less
than some extremely small value. This apparent paradox was resolved by A. Vainshtein in1972.
8He found that Dr. G’s theory contains a distance scale
rV=/parenleftBigg
GM
m4
G/parenrightBigg1
5
in the gravitational field around a body of mass M. The helicity 0 degree of freedom
becomes effective only on the distance scales r/greatermuchrV. Inside the Vainshtein radius rV, the
gravitational field is the same as in Einstein’s theory and experiments cannot distinguishbetween Einstein’s and Dr. G’s theories. With the current astrophysical bound m
G/lessmuch
(1024cm)−1andMthe mass of the sun, rVcomes out to be much larger than the size
of the solar system. In other words, the apparent paradox arose because of an interchangeof limits: We can take either the characteristic distance of the measurement r
obs(the radius
of the sun in the deflection of starlight) or the Vainshtein radius rVto infinity first.
So all is well: Dr. G’s theory is consistent with current measurements provided that
he takes mGsmall enough. What he is not allowed to do is use the one graviton exchange
approximation. Instead, he should solve the massive analog of Einstein’s field equation (4)around a massive body such as the sun, as Vainshtein did. This is equivalent to expandingto all orders in the graviton field hand resumming: In Feynman diagram language we
8A. I. Vainshtein, Phys. Lett. 39B:393, 1972; see also C. Deffayet, G. Dvali, G. Gabadadze, and A. I. Vainshtein,
Phys. Rev. D65:044026, 2002.
VIII.1. Gravity as a Field Theory | 441
q
k1 k2p1 p2
Figure VIII.1.2
have an infinite number of diagrams corresponding to the sun emitting 1, 2, 3, ... ,∞
gravitons respectively. The paradox is formally resolved by noting that the higher ordersare increasingly singular as m
G→0.
The gravity of light
At this point, you are ready to do perturbative quantum gravity: You have the graviton
propagator (12), and you can read off the interaction between gravitons from the detailedversion of (7) and the interaction between the graviton and any other field from the term−
1
2hμνTμν. The only trouble is that you might “drown in a sea of indices” if you don’t watch
out, as I have already warned you.
I know of one calculation (in fact one of my favorites in theoretical physics) in which
we can beat the indices down easily. An interesting question: Einstein said that lightis deflected by a massive object, but is light deflected gravitationally by light? T olman,Ehrenfest, and Podolsky discovered that in the weak field limit two light beams movingin the same direction do not interact gravitationally, but two light beams moving in theopposite directions do. Surprising, eh?
The scattering of two photons k
1+k2→p1+p2via the exchange of a graviton is given
by the Feynman diagram in figure VIII.1.2, with the momentum transfer q≡p1−k1, plus
another diagram with p1andp2interchanged. The Feynman rule for coupling a graviton
to two photons can be read off from
hμνTμν=−hμν(FμλFνλ−1
4ημνFρλFρλ)
but all we need is that the interaction involve two powers of spacetime derivatives ∂acting
on the electromagnetic potential Aμso that the graviton-photon-photon vertex involves
2 powers of momenta, one from each photon. Hence the scattering amplitude (with all
Lorentz indices suppressed) has the schematic form ∼(k1p1)D(k 2p2). The η’s in the
graviton propagator Dtie the indices on (k1p1)and(k2p2)together. (We have suppressed
the polarization vectors of the photons, imagining that they are to be averaged over in
the amplitude squared.) Referring to (12), we see that the amplitude is the sum of three
terms such as ∼(k1.p1)(k 2.p2)/q2,∼(k1.k2)(p 1.p2)/q2, and ∼(k1.p2)(k 2.p1)/q2.
Since according to Fourier the long distance part of the interaction potential is given by
442 | VIII. Gravity and Beyond
the small qbehavior of the scattering amplitude, we need only evaluate these terms in the
limitq→0. We can throw almost everything away! For example,
k1.p1→k1.k1=0,k1.p2=k1.(k1+k2−p1)→k1.k2
Just imagining contracting all those indices in our heads is good enough: We obtain the
amplitude ∼(k1.k2)(p1.p2)/q2.
Ifk1andk2point in the same direction, k1.k2∝k1.k1=0. Two photons moving in the
same direction do not interact gravitationally.
Of course, this result is not of any practical importance since electromagnetic effects are
far more important, but this is not an engineering text. In appendix 1 I give an alternativederivation of this amusing result.
Kaluza-Klein compactification
You have probably read about how excited Einstein was when he heard of the proposal ofKaluza and of Klein to extend the dimension of spacetime to 5 and thus unify electromag-netism and gravity. The 5th dimension is supposed to be compactified into a tiny circle ofradius afar smaller than what experimentalists can see; in other words, x
5is an angular
variable with x5=x5+2πa. You have surely heard that string theory, at least in some ver-
sion, is based on the Kaluza-Klein idea. Strings live in 10-dimensional spacetime, with 6of the dimensions compactified.
I can now show you how the Kaluza-Klein mechanism works. Start with the action
S=1
16πG 5/integraldisplay
d5x/radicalbig
−g 5R5 (15)
in 5-dimensional spacetime. The subscript 5 serves to indicate the 5-dimensional quanti-
ties. We denote the 5-dimensional metric by gABwith the indices AandBrunning over
0, 1, 2, 3, 5.
Assume that gABdoes not depend on x5. Plug into S, integrate over x5, and compute
the effective 4-dimensional action. Since R5and the 4-dimensional scalar curvature R
both involve two powers of ∂andgAB contains gμν, we must have (exercise VIII.1.5)
R5=R+.... Thus, (15) contains the Einstein-Hilbert action with Newton’s gravitational
constant G∼G5/a.
What else do we get? We don’t even have to work through the arithmetic. We can
argue by symmetry. Under the 5-dimensional coordinate transformation xA→x/primeA=
xA+εA(x), we have [see (8)] h/prime
AB=hAB−∂AεB−∂BεA. Let us choose εμ=0 andε5(x) to
be independent of x5: We go around and rotate each of the tiny circles attached to every point
in our spacetime a tiny bit. Well, we have h/prime
μν=hμνandh/prime
55=h55, buth/prime
μ5=hμ5−∂με5.
But if we give the Lorentz 4-vector hμ5and 4-scalar ε5new names, call them Aμand/Lambda1,
this just says A/prime
μ=Aμ−∂μ/Lambda1, the usual electromagnetic gauge transformation!
Since we know that the 5-dimensional action (15) is invariant under xA→x/primeA=xA+
εA(x), the resulting 4-dimensional action must be invariant under Aμ→A/prime
μ=Aμ−∂μ/Lambda1
VIII.1. Gravity as a Field Theory | 443
and hence must contain the Maxwell action. Note once again the power of symmetry
considerations. No need to do tedious calculations.
Electromagnetism comes out of gravity!
Differential geometry of Riemannian manifolds
I hinted earlier at a deep connection between general coordinate transformation and gaugetransformation. Let us flesh this out by looking at differential geometry and gravity. Forthis sketch we will consider locally Euclidean (rather than Minkowskian) spaces.
The differential geometry of Riemannian manifolds can be elegantly summarized in
the language of differential forms. Consider a Riemannian manifold (such as a sphere)with the metric g
μν(x). Locally, the manifold is Euclidean by definition, which means
gμν(x)=ea
μ(x)δabeb
ν(x) (16)
where the matrix e(x) may be thought of as a similarity transformation that diagonalizes
gμνand scales it to the unit matrix. Thus, for a D-dimensional manifold there exist D
“world vectors” ea
μ(x) obviously dependent on xand labeled by the index a=1, 2, ... ,D.
The functions ea
μ(x) are known as “vielbeins” (meaning “many legs” in German, vierbeins
=four legs for D=4, dreibeins =three legs for D=3, and so on.) In some sense, the
vielbeins can be thought as the “square root” of the metric.
Let us clarify by a simple example. The familiar two-sphere (of unit radius) has the line
element9ds2=dθ2+sin2θdϕ2. From the metric (g θθ=1,gϕϕ=sin2θ)we can read off
e1
θ=1 and e2
ϕ=sinθ(all other components are zero). We are invited to define D1-forms
ea=ea
μdxμ. (In our example, e1=dθ,e2=sinθdϕ .)
On a curved manifold, when we parallel transport a vector, the vector changes when
expressed in terms of the locally Euclidean coordinate frame. (This is just the familiarstatement that on a curved manifold such as the surface of the earth the notion of a vectorpointing straight north is a local concept: When we move infinitesimally away keeping our“north vector” pointing in the same direction, it will end up being infinitesimally rotatedaway from the “north vector” defined at the point we have just moved to.) This infinitesimalrotation of the vielbeins is described by
dea=−ωabeb(17)
Note that since ωgenerates an infinitesimal rotation it is an antisymmetric matrix: ωab=
−ωba. Since deais a 2-form, ωis a 1-form, known as the connection: It “connects” the
locally Euclidean frames at nearby points. (Since the indices a,b, etc. are associated with
the Euclidean metric δabwe do not have to distinguish between upper and lower indices.
When we do write upper or lower indices a,b, etc. it is for typographical convenience.) In
9Note that this represents the square of an infinitesimal distance element and not an area element, and so a
quantity such as dθ2is literally the square of dθand not the wedge product dθdθ (of chapter IV .4), which would
have been identically zero.
444 | VIII. Gravity and Beyond
the simple example of the sphere, de1=0 and de2=cosθd θd ϕ and so the connection
has only one nonvanishing component ω12=−ω21=−cos θd ϕ .
At any point, we are free to rotate the vielbeins: If you use the vielbeins ea
μI am free to use
some other vielbeins e/primea
μinstead, as long as mine are related to yours by a rotation ea
μ(x)=
Oa
b(x)e/primeb
μ(x). [You can check that gμν(x)=ea
μ(x)δabeb
ν(x)=e/primea
μ(x)δabe/primeb
ν(x) ifOTO=1.]
The connection ω/primeis defined by de/primea=−ω/primeabe/primeb. You can readily work out that (suppressing
indices)
ω=Oω/primeOT−(dO)OT(18)
The local curvature of the manifold is a measure of how the connection varies from
point to point. We would like the curvature to be invariant under the local rotation O(or at
least to transform as a tensor so that by contracting it with vectors we can form a scalar).The desired object is the 2-form R
ab=dωab+ωacωcb. You can check that R=OR/primeOT.
(For the sphere, R12=dω12+ω1cωc2=sinθd θd ϕ .)Written out in components, Rab=
Rab
μνdxμdxν. I leave it to you to verify that Rab
μνeλ
aeσ
bis the usual Riemann curvature tensor
Rλσ
μν, where eλ
ais the inverse of the matrix ea
λ. In particular, Rab
μνeμ
aeν
bis the scalar curvature,
which in our convention works out to be +1 for the sphere.
Thus, Riemannian geometry can be elegantly summarized by the two statements (again
suppressing indices)
de+ωe=0 (19)
and
R=dω+ω2(20)
Look familiar? You should be struck by the similarity between (20) and the expression
for the field strength in nonabelian gauge theories F=dA+A2. Note ωtransforms [see
(18)] exactly the same way as the gauge potential A. But one nagging difference, namely
the lack of an analog of ein gauge theory, has long bothered some theoretical physicists
(but is shrugged off by most as inconsequential). Also, note that Einstein theory is linearinRwhile Y ang-Mills theory is quadratic in F.
Gravity and Y ang-Mills
We can make the connection between gravity and Y ang-Mills theory more explicit by
looking at the derivative of a vector field. Y ang-Mills theory was born of the requirement
that a field ϕand its derivative ∂μϕtransform in the same way under a spacetime-
dependent internal symmetry transformation (IV .5.1). In Einstein gravity a vector field
Wμ(x) transforms as W/primeμ(x/prime)=Sμ
ν(x)Wν(x) withSμ
ν(x)=∂x/primeμ/∂xν. Since the matrix S
depends on the spacetime coordinate x, we see that ∂λWμcould not possibly transform
like a tensor with one upper and one lower index, as we would like naively just by lookingat indices. We would have to introduce a covariant derivative. Not surprisingly, this closely
VIII.1. Gravity as a Field Theory | 445
parallels the discussion in chapter IV .5. Historically, Y ang and Mills were inspired by
Einstein gravity.
Using the chain rule and the product rule, we have
∂/prime
λW/primeμ(x/prime)=∂W/primeμ(x/prime)
∂x/primeλ=∂xρ
∂x/primeλ∂
∂xρ[Sμ
ν(x)Wν(x)]=(S−1)ρ
λSμ
ν∂ρWν+[(S−1)ρ
λ∂ρSμ
ν]Wν(21)
Were the second term in (21), which comes from differentiating S, not there, the naive
guess, that ∂λWμtransforms like a tensor, would be valid. The fact that the transformation
Svaries from place to place has negated the naive guess.
What is happening is quite clear: as the vector Wvaries from a given point to a neighbor-
ing point, the coordinate axes that define the components of Walso change. This suggests
that we could define a more suitable derivative, called the covariant derivative and writtenasD
λWμ, to take this effect into account, so that DλWμwould indeed transform like a
tensor. Exactly as in Y ang-Mills theory (IV .5.1), we have to add an extra term to knock outthe second term in (21).
Just the way the indices hang together immediately suggests the correct construction.
The factor (S
−1)ρ
λ∂ρSμ
νin the unwanted second term in (21) has one upper index and two
lower indices, so we need an object with this set of indices. Lo, the Riemann-Christoffelsymbol /Gamma1
μ
λνin (3) (and introduced in chapter I.11) fits the bill perfectly. I will let you have
the fun of verifying that the covariant derivative defined by
DλWμ≡∂λWμ+/Gamma1μ
λνWν(22)
indeed transforms like a tensor (note that /Gamma1was normalized correctly for this purpose).
I end with a technical remark about the coupling of gravity to spin1
2fields. First,
we of course have to Wick rotate so that the vierbein ea
μerects a locally Minkowskian
rather than a Euclidean coordinate frame. The indices a,b, etc. are now contracted with
the Minkowskian metric ηab. The slight subtlety is that the Dirac gamma matrices γa
are associated with the Lorentz rotation of the vierbein ea
μ(x)=Oa
b(x)e/primeb
μ(x/prime)and thus
carry the Lorentz index arather than the “world” index μ. Similarly, the Dirac spinor
ψ(x) is defined relative to the local Lorentz frame specified by the vierbein, and thus
its covariant derivative has to be defined in terms of the connection ωrather than the
symbol /Gamma1. Hence the flat space Dirac action/integraltext
d4x¯ψ(iγμ∂μ−m)ψ must be general-
ized to/integraltext
d4x√−g¯ψ(iγaηabebμDμ−m)ψ , where the covariant derivative Dμψ=∂μψ−
i
4ωμabσabψexpresses the rotation of the local Lorentz frame as we move from a point xto
a neighboring point. In contrast to the action for integer spin fields in curved spacetime
(see chapter I.11), the Dirac action in curved spacetime involves the vierbein explicitly.
Appendix 1: Light on light again
The stress-energy tensor Tμνof a light beam moving in the x-direction has four nonzero components: the
energy density T00of course, then T0x=T00since photons carry the same energy and momentum, next
Tx0=T0xby symmetry, and finally Txx=T00since the stress-energy tensor of the electromagnetic field is
traceless (chapter I.11). Without having to solve Einstein’s equations in the weak field limit (13) explicitly weknow immediately that h
00=h0x=hx0=hxx≡h. The metric around the light beam is given by g00=1+h,
446 | VIII. Gravity and Beyond
g0x=gx0=−h, and gxx=− 1+h(and of course gyy=gzz=− 1 plus a bunch of vanishing components).
Consider a photon moving parallel to the light beam. Its worldline is determined by (recall chapter I.11)
d2xρ
dζ2=−/Gamma1ρ
μνdxμ
dζdxν
dζ
Let’s calculate d2y/dζ2andd2z/dζ2with(dy/dζ) ,(dz/dζ) /lessmuch(dt/dζ) ,(dx/dζ). Using (3) we find (with μ,ν
restricted to 0, x)
d2y
dζ2=1
2(∂νgyμ+∂μgyν−∂ygμν)dxμ
dζdxν
dζ
=−1
2(∂yh)/bracketleftbigg
(dt
dζ)2+(dx
dζ)2−2dt
dζdx
dζ/bracketrightbigg
=−1
2(∂yh)(dt
dζ−dx
dζ)2
For a photon moving in the same direction as the light beam dt=dxandd2y/dζ2=d2z/dζ2=0. We have once
again derived the T olman-Ehrenfest-Podolsky effect. Note we never had to solve for h.
Incidentally, if you are a bit unsure of dt=dx, the condition ds=0 for a light beam moving in the x-direction
amounts to (1+h)dt2−2hdtdx −(1−h)dx2=0. Upon division by dt2we obtain −(1+h)+2hv+(1−h)v2=
0, with v≡dx/dt . The quadratic equation has two roots v=∓(1±h)/(1−h). The negative root gives v=1,
and thus for a photon moving in the same direction as the light beam dx/dt =1. In contrast, the positive root
v=−(1+h)/(1−h)describes a photon moving in the opposite direction.
Appendix 2: The helicity structure of gravity
T o gain a deeper understanding of the difference between Einstein’s and Dr. G’s theories let us look at the
helicity structure of the interaction in the two cases. T o warm up, consider the interaction between two conserved
currents due to the exchange of a spin 1 particle of momentum kand mass m:Jμ
(1)J(2)μ=J0
(1)J0
(2)−Ji
(1)Ji
(2). Use
current conservation kμJμ=0 to eliminate J0=kiJi/ω(withω≡k0). We obtain (kikj/ω2−δij)Ji
(1)Jj
(2). Let/vectork
point in the 3rd direction and use /vectork2=ω2−m2to write this as −[(m2/ω2)J3
(1)J3
(2)+J1
(1)J1
(2)+J2
(1)J2
(2)]. We see
that as m→0 the longitudinal component of the current J3indeed decouples as explained in chapter II.7 and
we obtain −1
2(J1+i2
(1)J1−i2
(2)+J1−i2
(1)J1+i2
(2)), showing explicitly that the photon has helicity ±1. (Obvious notation:
J1+i2≡J1+iJ2etc.)
Onward to gravity. Consider the interaction Tμν
(1)T(2)μν−ξT(1)T(2), where ξ=1
2for Einstein and1
3for Dr. G.
For ease of writing I will now omit the subscripts (1) and (2). Conservation kμTμν=0 allows us to eliminate
T0i=kjTji/ωandT00=kjklTjl/ω2. Again taking /vectorkto point in the 3rd direction we obtain the mess
/parenleftbiggm
ω/parenrightbigg4
T33T33+2/parenleftbiggm
ω/parenrightbigg2
(T13T13+T23T23)+T11T11+T22T22+2T12T12
−ξ/bracketleftBigg/parenleftbiggm
ω/parenrightbigg2
T33+T11+T22/bracketrightBigg/bracketleftBigg/parenleftbiggm
ω/parenrightbigg2
T33+T11+T22/bracketrightBigg
which simplifies in the limit m→0t o
T11T11+T22T22+2T12T12−ξ(T11+T22)(T11+T22)
In Einstein’s theory, ξ=1
2and this becomes
1
2(T11−T22)(T11−T22)+2T12T12
which lo and behold is equal to1
2(T1+i2,1+i2T1−i2,1−i2+T1−i2,1−i2T1+i2,1+i2), showing that indeed the graviton
carries helicity ±2. In Dr. G’s theory, this would not be the case.
VIII.1. Gravity as a Field Theory | 447
Exercises
VIII.1.1 Work out Tμνfor a scalar field. Draw the Feynman diagram for the contribution of one-graviton exchange
to the scattering of two scalar mesons. Calculate the amplitude and extract the interaction energy betweentwo mesons sitting at rest, thus deriving Newton’s law of gravity.
VIII.1.2 Work out T
μνfor the Y ang-Mills field.
VIII.1.3 Show that if hμνdoes not satisfy the harmonic gauge, we can always make a gauge transformation with
ενdetermined by ∂2εν=∂μhμ
ν−1
2∂νhλλso that it does. All of this should be conceptually familiar from
your study of electromagnetism.
VIII.1.4 Count the number of degrees of polarization of a graviton. [Hint: Consider a plane wave hμν(x)=
hμν(k)eikxjust because it is a bit easier to work in momentum space. A symmetric tensor has 10
components and the harmonic gauge kμhμ
ν=1
2kνhλλimposes 4 conditions. Oops, we are left with 6
degrees of freedom. What is going on?] [Hint: You can make a further gauge transformation and stillstay in the harmonic gauge. The graviton should have only 2 degrees of polarization.]
VIII.1.5 The Kaluza-Klein result that we argued by symmetry considerations can of course be derived explicitly.
Let me sketch the calculation for you. Consider the metric
ds2=gμνdxμdxν−a2[dθ+Aμ(x)dxμ]2
where θdenotes an angular variable 0 ≤θ< 2π. With Aμ=0, this is just the metric of a curved spacetime,
which has a circle of radius aattached at every point. The transformation θ→θ+/Lambda1(x) leaves dsinvariant
provided that we also transform Aμ(x)→Aμ(x)−∂μ/Lambda1(x) . Calculate the 5-dimensional scalar curvature
R5and show that R5=R4−1
4a2FμνFμν. Except for the precise coefficient1
4this result follows entirely
from symmetry considerations and from the fact that R5involves two derivatives on the 5-dimensional
metric, as explained in the text. After some suitable rescaling this is the usual action for gravity pluselectromagnetism. Note that the 5-dimensional metric has the explicit form
g5
AB=/parenleftBigggμν−a2AμAν−a2Aμ
−a2Aν −a2/parenrightBigg
(23)
VIII.1.6 Generalize the Kaluza-Klein construction by replacing the circles by higher dimensional spheres. Show
that Y ang-Mills fields emerge.
VIII.1.7 Starting with the connection 1-form ω12=−cos θdϕ for the sphere, show that the scalar curvature is a
constant independent of θandϕ.
VIII.1.8 The vielbeins for a spacetime with Minkowski metric is defined by gμν(x)=ea
μ(x)ηabeb
ν(x), where the
Minkowski metric ηabreplaces the Euclidean metric δab. The indices aandbare to be contracted with
ηab. For example, Rab=dωab+ωacηcdωdb. Show that everything goes through as expected.
VIII.2The Cosmological Constant Problem
and the Cosmic Coincidence Problems
The force that knows too much
The word paradox has been debased by loose usage in the physics literature. A real paradox
should involve a major and clear-cut discrepancy between theoretical expectation andexperimental measurement. The ultraviolet catastrophe, for example, is a paradox, theresolution of which around the dawn of the twentieth century ushered in quantum physics.I now come to the most egregious paradox of present day physics.
The electromagnetic force knows about the particles carrying charge, and the strong
force knows about the particles carrying color. And the gravitational force? It knowseverybody! More precisely, anybody carrying energy and momentum.
Within a particle physics frame of mind, which is the only frame of mind we have
in exploring the fundamental structure of physics, the graviton can be regarded as justanother particle. Indeed, given that a massless spin 2 particle couples to the stress-energytensor, one can reconstruct Einstein’s theory.
Nevertheless, there is an uncomfortable feel to this whole picture. Gravity has to do with
the curvature of spacetime, the arena in which all fields and particles live. The graviton isnot just another particle.
This in essence is the root origin
1of the paradox of the cosmological constant. The
graviton is not just another particle—it knows too much!
The cosmological constant
In the absence of gravity, the addition of a constant /Lambda1to the Lagrangian L→L−/Lambda1has
no effect whatsoever. In classical physics the Euler-Lagrange equations of motion depend
1For more along this line, see A. Zee, hep-th/0805.2183 in Proceedings of the Conference in Honor of C. N. Y ang’s
85th Birthday, World Scientific, Singapore 2008, p. 131.
VIII.2. Cosmic Coincidence Problem | 449
only on the variation of the Lagrangian. In quantum field theory we have to evaluate the
functional integral Z=/integraltext
Dϕei/integraltext
d4xL(x), which upon the inclusion of /Lambda1merely acquires
a multiplicative factor. As we have seen repeatedly, a multiplicative factor in Zdoes not
enter into the calculation of Green’s function and scattering amplitudes.
Gravity, however, knows about /Lambda1. Physically, the inclusion of /Lambda1corresponds to a shift
in the Hamiltonian H→H+/integraltext
d3x/Lambda1. Thus, the “cosmological constant” /Lambda1describes a
constant energy or mass per unit volume permeating the universe, and of course gravityknows about it.
More technically, the term in the action −/integraltext
d
4x/Lambda1 is not invariant under a coordinate
transformation x→x/prime(x). In the presence of gravity, general coordinate invariance re-
quires that the term −/integraltext
d4x/Lambda1 in the action Sbe modified to −/integraltext
d4x√g/Lambda1, as I explained
way back in chapter I.11. Thus, the gravitational field gμνknows about /Lambda1, the infamous
cosmological constant introduced by Einstein and lamented by him as his biggest mistake.This often quoted lament is itself a mistake. The introduction of the cosmological constantis not a mistake: It should be there.
Symmetry breaking generates vacuum energy
In our discussion on spontaneous symmetry breaking, we repeatedly ignored an additivetermμ
4/4λthat appears in L.
Particle physics is built on a series of spontaneous symmetry breaking. As the universe
cools, grand unified symmetry is spontaneously broken, followed by electroweak symme-try breaking, then chiral symmetry breaking, just to mention a few that we have discussed.At every stage a term like μ
4/4λappears in the Lagrangian, and gravity duly takes note.
How large do we expect the cosmological constant /Lambda1to be? As we will see, for our
purposes the roughest order of magnitude estimate suffices. Let us take λto be of order
1. As for μ, for the three kinds of symmetry breaking I just mentioned, μis of order 1017,
102, and 1 Gev, respectively. We thus expect the cosmological constant /Lambda1to be roughly
μ4=μ/(μ−1)3, where the last form of writing μ4reminds us that /Lambda1is a mass or energy
density: An energy of order μpacked into a cube of size μ−1. But this is outrageous even if
we take the smallest value for μ: We know that the universe is not permeated with a mass
density of the order of 1 Gev in every cube of size 1 (Gev)−1.
We don’t have to put in actual numbers to see that there is a humongous discrepancy be-
tween theoretical expectation and observational reality. If you want numbers, the currentobservational bound on the cosmological constant is <∼(10
−3ev)4. With the grand unifi-
cation energy scale, we are off by (17+9+3)×4=116 orders of magnitude. This is the
mother of all discrepancies!
With the Planck mass MPl∼1019Gev the natural scale of gravity, we would expect
/Lambda1∼M4
Plif it is of gravitational origin. We are then off by 124 orders of magnitude. We
are not talking about the crummy calculation of some pitiful theorist not fitting someexperimental curve by a factor of 2.
450 | VIII. Gravity and Beyond
We can imagine the universe starting out with a negative cosmological constant, fined
tuned to cancel the cosmological constant generated by the various episodes of sponta-neous symmetry breaking. Or there must be a dynamical mechanism that adjusts thecosmological constant to zero.
Notice I say zero, because the cosmological constant problem is basically an enormous
mismatch between the units natural to particle physics and natural to cosmology. Measuredin units of Gev
4the cosmological constant is so incredibly tiny that particle physicists
have traditionally assumed that it must be zero and have looked in vain for a plausiblemechanism to drive it to zero. One of the disappointments of string theory is its inabilityto resolve the cosmological constant problem. As of the writing of this chapter around theturn of the millennium, the brane world scenarios (chapter I.6) have generated a great dealof excitement by offering a glimmer of a hope. Roughly, the idea is that the gravitationaldynamics of the larger space that our universe is embedded in may cancel the effect of thecosmological constant.
Cosmic coincidence
But Nature has a big surprise for us. While theorists racked their brains trying to come upwith a convincing argument that /Lambda1=0, observational cosmologists steadily refined their
measurements and discovered dark energy. The “cleanest” explanation of dark energy byfar is that it represents the cosmological constant. Assuming that this is the case (andwho knows?), the upper bound on the cosmological constant would be changed to anapproximate equality
/Lambda1∼(10−3ev)4!!! (1)
The cosmological constant paradox deepens. Theoretically, it is easier to explain why some
quantity is mathematically 0 than why it happens to be ∼10−124in the units natural (?) to
the problem.
T o make things worse, (10−3ev)4happens to be the same order of magnitude as the
present matter density of the universe ρM. More precisely, dark energy accounts for ∼74%
of the mass content of the universe, dark matter for ∼22%, and ordinary matter for ∼4%.
First, the ordinary matter we know and love is reduced to an almost negligibly smallcomponent of the universe. Second, why should ρ
Mbe comparable to /Lambda1to within a factor
of 3? This is sometimes referred to as the cosmic coincidence problem.
Now the cosmological constant /Lambda1is, within our present understanding, a parameter in
the Lagrangian. On the other hand, since most of the mass density of the universe residesin rest mass, as the universe expands ρ
M(t)decreases as [1 /R(t)]3, where R(t) denotes
the scale size of the universe.2In the far past, ρMwas much larger than /Lambda1, and in the
2For an easy introduction to cosmology, see A. Zee, Unity of Forces in the Universe , vol. II, chap. 10.
VIII.2. Cosmic Coincidence Problem | 451
far future, it will be much smaller. It just so happens that, in this particular epoch of the
universe, when you and I are around, ρM∼/Lambda1. Or to be less anthropocentric, the epoch
when ρM∼/Lambda1happens to be when galaxy formation has been largely completed.
Very bizarre!In their desperation, some theorists have even been driven to invoke anthropic
selection.
3
3For a recent review, see A. Vilenkin, hep-th/0106083.
VIII.3Effective Field Theory Approach
to Understanding Nature
Low energy manifestation
The pioneers of quantum field theory, Dirac for example, tended to regard field theory
as a fundamental description of Nature, complete in itself. As I have mentioned severaltimes, in the 1950s, after the success of quantum electrodynamics many leading particlephysicists rejected quantum field theory as incapable of dealing with the strong and weakinteractions, not to mention gravity. Then came the great triumph of field theory in the early1970s. But after particle physicists retrieved field theory from the dust bin of theoreticalphysics, they realized that the field theories they were studying might be “merely” the lowenergy manifestation of a deeper structure, a structure first identified as a grand unifiedtheory and later as a string theory. Thus was developed an outlook known as the effectivefield theory approach, pace Dirac.
The general idea is that we can use field theory to say something about physics at low
energies or equivalently long distances even if we don’t know anything about the ultimatetheory, be it a theory built on strings or some as yet undreamed of structure. An importantconsequence of this paradigm shift was that nonrenormalizable field theories becameacceptable. I will illuminate these remarks with specific examples.
The emergence of this effective field theory philosophy, championed especially by Wil-
son, marks another example of cross fertilization between condensed matter and particlephysics. T oward the late 1960s, Wilson and others developed a powerful effective field the-ory approach to understanding critical phenomena, culminating in his Nobel Prize. Thesituation in condensed matter physics is in many ways the opposite of that in particle phys-ics at least as particle physics was understood in the 1960s. Condensed matter physicistsknow the short distance physics, namely the quantum mechanics of electrons and ions.But it certainly doesn’t help in most cases to write down the Schr ¨odinger equation for the
electrons and ions. Rather, what one would like to have is an effective description of howa system would respond when probed at low frequency and small wave vector. A strikingexample is the effective theory of the quantum Hall fluid as described in chapter VI.2: The
VIII.3. Effective Field Theory | 453
relevant degree of freedom is a gauge field, certainly a far cry from the underlying elec-
tron. As in the σmodel description (chapter VI.4) of quantum chromodynamics, it is fair
to say that without experimental guidance theorists would have a terribly hard time de-ciding what the relevant low energy long distance degrees of freedom might be. You haveseen numerous other examples in condensed matter physics, from the Landau-Ginzburgtheory of superconductivity to Peierls instability.
The threshold of ignorance
In our discussion of renormalization, I espouse the philosophy that a quantum fieldtheory provides an effective description of physics up to a certain energy scale /Lambda1,a
threshold of ignorance beyond which physics not included in the theory comes intoplay. In a nonrenormalizable theory, various physical quantities that we might wish tocalculate will come out dependent on /Lambda1, thus indicating that the physics at or beyond
the scale /Lambda1is essential for understanding the low energy physics we are interested in.
Nonrenormalizable theories suffer from not being totally predictive, but nevertheless theymay be useful. After all, the Fermi theory of the weak interaction described experimentsand even foretold its own demise.
In a renormalizable theory, various physical quantities come out independent of /Lambda1,
provided that the calculated results are expressed in terms of physical coupling constantsand masses, rather than in terms of some not particularly meaningful bare couplingconstants and masses. Low energy physics is not sensitive to what happens at highenergies, and we are able to parametrize our ignorance of high energy physics in terms ofa few physical constants.
From the late 1960s to the 1970s, one main thrust of fundamental physics was to
classify and study renormalizable theories. As we know, this program was “more thanspectacularly successful.” It allowed us to pin down the theory of the strong, the weak, andthe electromagnetic interactions.
Renormalization group flow and dimensional analysis
The effective field theory philosophy is intrinsically tied to renormalization group flow.In a given field theory, as we flow toward low energies, some couplings may tend to zerowhile others do not (and if they tend to infinity as in QCD, then we are unable to figureout the effective theory without experimental input). Thus, the first step is to calculate therenormalization group flow. A simple example is given in exercise VIII.3.1.
In many cases, we can simply use dimensional analysis. As I explained in our earlier
discussion on renormalization theory, couplings with negative dimensions of mass are
not important at low energies. T o be specific, suppose we add a gϕ
6term to a λϕ4theory.
The coupling ghas the dimension of inverse mass squared. Let us define M2≡1/g.A t
low energies, the effect of the gϕ6term is suppressed by (E/M)2.
454 | VIII. Gravity and Beyond
How do we understand Schwinger’s spectacular calculation of the anomalous magnetic
moment of the electron in the effective field theory philosophy?
Let me first tell the traditional (i.e., pre-Wilsonian) version of the story. A student could
have asked, “Professor Schwinger, why didn’t you include the term (1/M)¯ψσμνψFμνin
the Lagrangian?”
The answer is that we better not. Otherwise, we would lose our prediction for the
anomalous magnetic moment; it would depend on M. Recall that [ ψ]=3
2and [Aμ]=1,
and hence ¯ψσμνψFμνhas mass dimension3
2+3
2+1+1=5>4. The requirement of
renormalizability, that the Lagrangian be restricted to contain operators of dimension 4 orless, provides the rationale for excluding this term.
Actually, the “real” punchline of my story is that Schwinger probably would not have
answered the question. When I took Schwinger’s field theory class, it was well knownamong the students that it was forbidden to ask questions. Schwinger would simplyignore any raised hands. There was no opportunity to ask questions after class either:As he uttered his last sentence of an invariably beautifully prepared lecture, he would sailmajestically out of the room. Dirac dealt with questions differently. I was too young to havewitnessed it, but the story goes that when a student asked, “Professor Dirac, I did not under-stand..., ” Dirac replied, “That is an assertion, not a question.”
The modern retelling of the magnetic moment story turns it around. We now regard the
Lagrangian of quantum electrodynamics as an effective Lagrangian which should includean infinite sequence of terms of ever higher dimensions, with coefficients parametrizingour threshold of ignorance. The physics of electrons and photons is now described by
L=¯ψ(iγμ(∂μ−ieAμ)−m)ψ−1
4FμνFμν+1
M¯ψσμνψFμν+...
Yes, the term (1/M)¯ψσμνψFμνis there, with some unknown Mhaving the dimension of
a mass. Schwinger’s result, that quantum fluctuations generate a term(α/2π)(1/2m
e)¯ψσμνψFμν, should then be interpreted as saying that the anomalous
magnetic moment of the electron is predicted to be [ (α/2π)(1/2me)+1/M]. The close
agreement of (α/2π)(1/2me)with the experimental value of the anomalous magnetic
moment can then be turned around to set a lower bound on M/greatermuch(4π/α)me.
Equivalently, Schwinger’s result predicts the anomalous magnetic moment of the elec-
tron if we have independent evidence that Mis much larger than [ (α/2π)(1/2me)]−1.I
want to emphasize that all of this makes total physical sense. For example, if you speculatethat the electron has some finite size a, then you would expect M∼1/a. The anomalous
magnetic moment calculation gives an upper bound for a, telling us that the electron must
be pointlike down to some small scale. Alternatively, we could have had independent evi-dence, from electron scattering for example, that ahas to be smaller than a certain length,
thus giving us a lower bound on M.
T o underscore this point, imagine that in 1948 we followed Schwinger and quickly
calculated the anomalous magnetic moment of the proton. We could literally have doneit in 3 seconds, since all we have to do is replace m
ebympin the Lagrangian, thus
obtaining (α/2π)(1/2mp)¯ψσμνψFμν, which would of course disagree resoundingly with
VIII.3. Effective Field Theory | 455
experiment. The disagreement tells us that we had not included all the relevant physics,
namely that the proton interacts strongly and is not pointlike. Indeed, we now know thatthe anomalous magnetic moment of the proton gets contributions from the anomalousmagnetic moments and the orbital motion of the quarks inside the proton.
Effective theory of proton decay
It may seem that with the effective field theory approach we lose some predictive power. Buteffective field theories can also be surprisingly predictive. Let me give a specific example.Suppose we had never heard of grand unified theory. All we know is the SU( 3)⊗SU( 2)⊗
U(1)theory. An experimentalist tells us that he is planning to see if the proton would decay.
Without the foggiest notion about what would cause the proton to decay we can still
write down a field theory to describe proton decay. The Lagrangian Lis to be constructed
out of quark qand lepton lfields and must satisfy the symmetries that we know. Three
quarks disappear, so we write down schematically qqq , but three spinors do not a Lorentz
scalar make. We have to include a lepton field and write qqql .
Since four fermion fields are involved, the terms qqql have mass dimension 6 and so in
Lthey have to appear as (1/M
2)qqql with some mass M, corresponding to the mass scale
of the physics responsible for proton decay. The experimental lower bound on the lifetimeof the proton sets a lower bound on M.
It is instructive to contrast this analysis with an (imagined) effective field theory analysis
of proton decay long before the concept of quarks was invented, say around 1950. We wouldconstruct an effective Lagrangian out of the available fields, namely the proton field p,
the electron field e, and the pion field π, and thus write down the dimension 4 operator
f¯pe
+π0with some dimensionless constant f. T o estimate f, we would naively compare
this operator with the one describing pion-nucleon coupling (chapter IV .2) g¯pnπ+in the
effective Lagrangian. Since f¯pe+π0violates isospin invariance, we might expect f∼αg,
namely the same order as gmultiplied by some measure of isospin breaking, say the fine
structure constant. But this would give an unacceptably short lifetime to the proton. Weare forced to set fto a ridiculously small number, which seems highly unnatural. Thus, at
least in hindsight, we can say that the extremely long lifetime of the proton almost pointsto the existence of quarks. The key, as we saw above, is to promote of the mass dimensionof the term in the effective Lagrangian responsible for proton decay from 4 to 6. (Can thecosmological constant puzzle be solved in the same way?)
Another way of saying this is that SU( 3)⊗SU( 2)⊗U(1)plus renormalizability predicts
one of the most striking facts of the universe, the stability of the proton. In contrast, theold pion-nucleon theory glaringly failed to explain this experimental fact.
In accordance with our philosophy, Lmust be invariant under SU( 3)⊗SU( 2)⊗U(1),
under which quark and lepton fields transform rather idiosyncratically, as we saw inchapter VII.5. T o construct Lwe have to sit down and list all Lorentz invariant SU( 3)⊗
SU( 2)⊗U(1)terms of the form qqql .
456 | VIII. Gravity and Beyond
Sitting down, we would find that, assuming only one family of quarks and leptons for
simplicity, there are only four terms we can write down for proton decay, which I list herefor the sake of completeness: (/tildewidel
LCqL)(uRCdR),(eRCuR)(/tildewideqLCqL),(/tildewidelLCqL)(/tildewideqLCqL), and
(eRCuR)(uRCdR). Here lL=/parenleftbigν
e/parenrightbig
LandqL=/parenleftbigu
d/parenrightbig
Ldenote the lepton and quark doublet
ofSU( 2)⊗U(1), the twiddle is defined by /tildewidelj=liεijwithSU( 2)indices i,j=1, 2 (see
appendix B), and Cdenotes the charge conjugation matrix. Color indices on the quark
fields are contracted in the only possible way. The effective Lagrangian is then given by thesum of these four terms, with four unknown coefficients.
The effective field theory tells us that all possible baryon number violating decay pro-
cesses can be determined in terms of four unknowns. We expect that these predictions willhold to an accuracy of order (M
W/M)2. (IfMWwere zero, SU( 3)⊗SU( 2)⊗U(1)would
be exact.)
Of course, we can increase our predictive power by making further assumptions. For
example, if we think that proton decay is mediated by a vector particle, as in a generic grandunified theory, then only the first two terms in the above list are allowed. In a specific grandunified theory, such as the SU( 5)theory, the two unknown coefficients are determined in
terms of the grand unified coupling and the mass of the Xboson.
T o appreciate the predictive power of the effective field theory approach, inspect the
list of the four possible operators. We can immediately predict that while proton decayviolates both baryon number Band lepton number L, it conserves the combination B−L.
I emphasize that this is not at all obvious before doing the analysis. Could you have toldthe experimentalist which of the two possible modes n→e
+π−orn→e−π+he should
expect? A priori, it could well be that B+Lis conserved.
Note that Fermi’s theory of the weak interaction would be called an effective field theory
these days. Of course, in contrast to proton decay, beta decay was actually seen, and theprediction from this sort of symmetry analysis, namely the existence of the neutrino, wastriumphantly confirmed.
Along the same line, we could construct an effective field theory of neutrino masses.
Surely one of the most exciting experimental discoveries in particle physics of recentyears was that neutrinos are not massless. Let us construct an SU( 2)⊗U(1)invariant
effective theory. Since ν
Lresides inside lL, without doing any detailed analysis we can see
that a dimension-5 operator is required: schematically lLlLcontains the desired neutrino
bilinear but it carries hypercharge Y/2=− 1; on the other hand, the Higgs doublet ϕ
carries hypercharge +1
2, and so the lowest dimensional operator we can form is of the
formllϕϕ with dimension3
2+3
2+1+1=5. Thus, the effective Lmust contain a term
(1/M)llϕϕ , with Mthe mass scale of the new physics responsible for the neutrino mass.
Thus, by dimensional analysis we can estimate mν∼m2
l/M, with mlsome typical charged
lepton mass. If we take mlto be the muon mass ∼102Mev and mν∼10−1ev, we find
M∼(102Mev)2/10−1(10−6Mev)=108Gev.
The philosophy of effective field theories valid up to a certain energy scale /Lambda1seems
so obvious by now that it is almost difficult to imagine that at one time many eminentphysicists demanded much more of quantum field theory: that it be fundamental up toarbitrarily high energy scales.
VIII.3. Effective Field Theory | 457
Indeed, we now regard all quantum field theories as effective field theory. For all we
know, spacetime on some short distance does consist of a lattice, and so the Y ang-Millsaction is but the leading term in an expansion of the Wilson lattice action. The Einstein-Hilbert Lagrangian, being nonrenormalizable, is a fortiori “merely" the leading term inan effective field theory
L=√−g(M4
/Lambda1+M2
PR+c1R2+c2RμνRμν+c3RμνσρRμνσρ+1
M2(d1R3+...)+...)
Herec1, 2, 3 andd1are dimensionless numbers presumably of order 1. The three terms
quadratic in the curvature involve four powers of derivatives versus the two powers inthe Einstein-Hilbert term, and hence their effects relative to the leading terms are sup-pressed by (E/M
P)2withEan energy scale characteristic of the process we are studying.
Thus, these so-called Weyl-Eddington terms could be safely ignored in any conceivableexperiment. [A technical aside: The Gauss-Bonnet theorem implies that the combination(R
2−4RμνRμν+RμνσρRμνσρ)is a total derivative, so c3can be effectively set to 0, but that
is besides the point here.] We have indicated only one representative dimension 6 termR
3(out of many). Its coefficient, in accordance with high school dimensional analysis, is
suppressed by two powers of some mass M.
What do we expect the mass scale Mto be? Suppose we live in a universe with only gravity
(and of course we don’t, actually) then once again, we could risk being presumptuous andtakeMto be the intrinsic mass scale of gravity, namely the Planck mass M
P, but we have
not yet recovered from our third-degree burn from supposing that M/Lambda1∼MP. If we could
ignore the cosmological constant problem for a moment, then the standard (but quitepossibly wrong!) consensus is that in a universe of pure gravity our theory of gravity is aneffective expansion in powers of (E/M
P)2.
Alternatively, we could treat Las the effective theory of gravity after we integrate out all
the matter degree of freedom. In that case, Mwould be of order me(imagine gravitons
coupled to an electron loop; see exercise VIII.3.5), or perhaps even mν(generated by a
neutrino loop).
Effective field theory of the blue sky
As another application of the effective field theory philosophy, consider the scattering of
electromagnetic waves on an electrically neutral spinless particle described by a scalar field
/Phi1. Since /Phi1is neutral, the lowest dimension gauge invariant term that can be added to
L=∂/Phi1†∂/Phi1+m2/Phi1†/Phi1+...is (1/M2)/Phi1†/Phi1FμνFμν. A factor of 1 /M2, with Msome mass
scale, has to be included with the dimension 1 +1+2+2=6 operator to bring the high
school dimension down to 4. The two powers of derivative in FμνFμνtell us immediately
that the amplitude for photon scattering on this neutral particle goes like M∝ω2, with
ωthe frequency of the electromagnetic wave. Thus we conclude that the scattering cross
section varies like σ(ω)∝ω4.
458 | VIII. Gravity and Beyond
We have arrived at Rayleigh’s celebrated explanation of the color of the sky. In passing
through the atmosphere red light scatters less than blue light on air molecules and hencethe sky is blue.
For application to spinless atoms or molecules, we can pass to the nonrelativistic limit
as described in chapter III.5, setting /Phi1=(1/√
2m)e−imtϕ, so that the effective Lagrangian
now reads
L=ϕ†i∂0ϕ−1
2m∂iϕ†∂iϕ+1
mM2ϕ†ϕ(c 1/vectorE2−c2/vectorB2)+...
In this case, since we understand the microscopic physics governing atoms and
molecules, we know perfectly well what the mass scale Mrepresents. The coupling of
a photon to an electrically neutral system such as an atom or a molecule must vanish likethe characteristic size dof the system, since as d→0 the positive and negative charges are
on top of each other, giving a vanishing net coupling to the photon. Rotational invarianceimplies that the coupling ∼/vectork./vectord. The scattering amplitude then goes like M∝(ωd)
2,
since the coupling has to act twice, once for the incoming photon and once for the out-going photon. (Note that by rotational invariance the expectation value of the operator /vectord
vanishes, but we are doing second order perturbation theory so that we have to evalu-ate the expectation value of a quantity quadratic in /vectord.)
1Squaring Mand invoking some
elementary quantum mechanics and dimensional analysis, we obtain the cross sectionσ(ω)∼d
6ω4.
Appendix: Reshuffling terms in effective field theory
The Lagrangian of an effective field theory consists of an infinite sequence of terms arranged in an orderly
progression of higher and higher mass dimension, constrained only by the assumed symmetries of the theory.In fact, some terms could be effectively eliminated. T o explain this, we focus on a toy example:
L=1
2(∂ϕ)2−λϕ4+1
M2(aϕ6+bϕ3∂2ϕ+c(∂2ϕ)2)+O/parenleftbigg1
M4/parenrightbigg
(1)
We are secretly dealing with the action and thus we freely integrate by parts. For arithmetical simplicity, we did
not include a mass term, so that to leading order in 1 /M the equation of motion reads simply ∂2ϕ=0. The three
possible dimension 6 terms are shown explicitly [we integrate by parts to get rid of the term ϕ2(∂ϕ)2].
Are we allowed to use the equation of motion to eliminate the two dimension 6 terms that are proportional
to∂2ϕ?
We know that we could make a field redefinition without changing the on shell amplitudes, so let us rede-
fineϕ→ϕ+(1/M2)F. Then1
2(∂ϕ)2→1
2(∂ϕ)2−(1/M2)F∂2ϕ+O(1/M4)andλϕ4→λ(ϕ4+(1/M2)ϕ3F+
O(1/M4)).S e tF=pϕ3+q∂2ϕ. We see that with an appropriate choice of pandqwe can cancel off bandc.
Notice that in the process we also change ato some other value.
The answer to the question is yes, but the naive statement that the equation of motion ∂2ϕ=0 empowers us
to simply set ∂2ϕto zero in the nonleading terms in the effective field theory is, legalistically speaking, incorrect,
or at least misleading. We see that we actually generated O(1/M4)terms and changed the ϕ6term. Thus, more
correctly, a field redefinition allows us to shuffle terms around and to higher order. The net effect, however, isthe same as if we trusted the naive statement and set ∂
2ϕto zero in the nonleading terms.
1For details, see, for example, J. J. Sakurai, Advanced Quantum Mechanics, Addison-Wesley, New York, 1967,
p. 47.
VIII.3. Effective Field Theory | 459
This procedure works for fermions also. As an example, consider the effective Lagrangian L=¯ψ(iγμ∂μ−
m)ψ+(1/M3)¯ψ(iγμ∂μ−m)ψ( ¯ψψ)+.... Then the field redefinition ψ→ψ−(1/2M3)ψ(¯ψψ) gets rid of the
dimension 7 term shown.
We could also apply what we just learned to the effective theory of gravity if without any understanding we
set the cosmological constant to zero. Also, use the Gauss-Bonnet theorem to get rid of the RμνσρRμνσρterm, so
that we have
L=√−g/parenleftbigg
M2
PR+c1R2+c2RμνRμν+1
M2(d1R3+...)+.../parenrightbigg
(2)
Make a field redefinition gμν→gμν+δgμνand use
δ/integraldisplay
d4x√−gR=−/integraldisplay
d4x√−g(Rμν−1
2gμνR)δgμν
Setδgμν=pRμν+qgμνR. Then we can cancel off c1andc2with a judicious choice of pandq. I emphasize that
this works only if we set the cosmological constant to zero without any ado.
Exercises
VIII.3.1 Consider
L=1
2/bracketleftBig
(∂ϕ 1)2+(∂ϕ2)2/bracketrightBig
−λ(ϕ4
1+ϕ4
2)−gϕ2
1ϕ2
2(3)
We have taken the O(2)theory from chapter I.10 and broken the symmetry explicitly. Work out the
renormalization group flow in the (λ−g)plane and draw your own conclusions.
VIII.3.2 Assuming the nonexistence of the right handed neutrino field νR(i.e., assuming the minimal particle
content of the standard model) write down all SU( 2)⊗U(1)invariant terms that violate lepton number
Lby 2 and hence construct an effective field theory of the neutrino mass. Of course, by constructing a
specific theory one can be much more predictive. Out of the product lLlLwe can form a Lorentz scalar
transforming as either a singlet or triplet under SU( 2). T ake the singlet case and construct a theory. [Hint:
For help, see A. Zee, Phys. Lett. 93B : p. 389, 1980.]
VIII.3.3 LetA,B,C,Ddenote four spin1
2fields and label their handedness by a subscript: γ5Ah=hAhwith
h=± 1. Thus, A+is right handed, A−left handed, and so on. Show that
(AhBh)(C−hD−h)=−1
2(AhγμD−h)(C−hγμBh) (4)
This is an example of a broad class of identities known as Fierz identities (some of which we will need in
discussing supersymmetry.) Argue that if proton decay proceeds in lowest order from the exchange of avector particle then only the terms (/tildewidel
LCqL)(uRCdR)and(eRCuR)(/tildewideqLCqL)are allowed in the Lagrangian.
VIII.3.4 Given the conclusion of the previous exercise show that the decay rate for the processes p→π++¯ν,
p→π0+e+,n→π0+¯ν, andn→π−+e+are proportional to each other, with the proportionality
factors determined by a single unknown constant [the ratio of the coefficients of (/tildewidelLCqL)(uRCdR)and
(eRCuR)(/tildewideqLCqL)].
For help on these last three exercises see S. Weinberg, Phys. Rev. Lett. 43: 1566, 1979; F. Wilczek and
A. Zee, ibid. p. 1571; H. A. Weldon and A. Zee, Nucl. Phys. B173: 269, 1980.
VIII.3.5 Imagine a mythical (and presumably impossible) race of physicists who only understand physics at
energies less than the electron mass me. They manage to write down the effective field theory for the one
particle they know, the photon,
L=−1
4FμνFμν+1
m4
e{a(FμνFμν)2+b(Fμν˜Fμν)2}+ ... (5)
460 | VIII. Gravity and Beyond
with ˜Fμν=1
2εμνρσFρσthe dual field strength as usual and aandbtwo dimensionless constants
presumably of order unity.(a) Show that Lrespects charge conjugation ( A→−Ain this context), parity, and time reversal, (and
of course gauge invariance.)
(b) Draw the Feynman diagrams that give rise to the two dimension 8 terms shown. The coefficients
aandbwere calculated by Euler and Kockel in 1935 and by Heisenberg and Euler in 1936, quite a
feat since they did not know about Feynman diagrams and any of the modern quantum field theoryset up.
(c) Explain why dimensional 6 terms are absent in L. [Hint: One possible term is ∂
λFμν∂λFμν.]
(d) Our mythical physicists do not know about the electron, but they are getting excited. They are going
to start doing photon-photon scattering experiments with a machine called LPC that could producephotons with energy greater than m
e. Discuss what they will see. Apply unitarity and the Cutkosky
rules.
VIII.3.6 Use the effective field theory approach to show that the scattering cross section of light on an electrically
neutral spin1
2particle (such as the neutron) goes like σ∝ω2to leading order, not ω4. Argue further that
the constant of proportionality can be fixed in terms of the magnetic moment μof the particle. [Historical
note: This result was first obtained in 1954 by F. Low ( Phys. Rev. 96: 1428) and by M. Gell-Mann and
Murph L. Goldberger (Phys. Rev. 96: 1433) using much more elaborate arguments.]
VIII.4 Supersymmetry: A Very Brief Introduction
Unifying bosons and fermions
Let me start with a few of the motivations for supersymmetry. (1) All experimentally known
symmetries relate bosons to bosons and fermions to fermions. We would like to have asymmetry, supersymmetry, relating bosons and fermions. (2) It is natural for fermions tobe massless (recall chapter VII.6), but not for bosons. Perhaps by pairing the Higgs fieldwith a fermion field we can resolve the hierarchy problem mentioned in chapter VII.6.(3) Recalling from chapter II.5 that fermions contribute negatively to the vacuum energy,you might be tempted to speculate that the cosmological constant problem could be solvedif we could get the fermion contribution to cancel the boson contribution.
Disappointingly, it has been more than 30 years
1since the conception of supersymmetry
(Golfand and Likhtman constructed the first supersymmetric field theory in 1971) anddirect experimental evidence is still lacking. All existing supersymmetric theories pairknown bosons with unknown fermions and known fermions with unknown bosons.Supersymmetry has to be broken at some mass scale Mbeyond the regime already explored
experimentally, but then (as explained in chapter VIII.2) we might expect a cosmologicalconstant of order M
4.
Be that as it may, supersymmetric field theories have many nice properties (hardly
surprising since the relevant symmetry is much larger). Supersymmetry has thus attracteda multitude of devotees. I give you here as brief an introduction to supersymmetry as I canwrite. In the spirit of a first exposure, I will avoid mentioning any subtleties and caveats,hoping that this brief introduction will be helpful to students before they tackle the tomesout there.
1For a fascinating account of the early history of supersymmetry, see G. Kane and M. Shifman, eds., The
Supersymmetric World: The Beginning of the Theory .
462 | VIII. Gravity and Beyond
Inventing supersymmetry
Suppose one day you woke up wanting to invent a field theory with a symmetry relating
bosons to fermions. The first thing you would need is the same number of fermionic andbosonic degrees of freedom. The simplest fermion field is the two-component Weyl spinorψ. You would now have one complex degree of freedom,
2so you would have to throw in
a complex scalar field ϕ. You could proceed by trial and error: Write down a Lagrangian
including all terms with dimension up to four and then adjust the various parameters inthe Lagrangian until the desired symmetry appears. For instance, you might adjust μin
the mass terms μ
2ϕ†ϕ+m(ψψ +¯ψ¯ψ) until the theory becomes more symmetrical so
that the boson and the fermion have the same mass.
If you were to try to play the game by using a Dirac spinor /Psi1and a complex scalar
ϕyou would be doomed to failure from the very start since there would be twice as
many fermionic degrees of freedom as bosonic degrees of freedom. I believe that thedevelopment of supersymmetry was very much retarded by the fact that until the early1970s most field theorists, having grown up with Dirac spinors, had little knowledge ofWeyl spinors. That was a hint that now is the time for you to get thoroughly familiar withthe dotted and undotted notation of appendix E. T o read this chapter, you need to be fluentwith that notation.
Supersymmetric algebra
It is perfectly feasible to construct this supersymmetric field theory, known as the Wess-Zumino model, by trial and error, but instead I will show you an elegant but more abstractapproach known as the superspace and superfield formalism, invented by Salam andStrathdee. We will have to develop a considerable amount of formal machinery. Everythingis very super here.
Write the supersymmetry generator taking us from ϕtoψ
αasQα(known as the
supercharge). The statement that Qαtransforms as a Weyl spinor means [ Jμν,Qα]=
−i(σμν)αβQβ, where Jμνdenotes the generators of the Lorentz group. Of course, since
Qαis independent of the spacetime coordinates [ Pμ,Qα]=0. From appendix E we denote
the conjugate of Qαby¯Q˙αand [Jμν,¯Q˙α]=−i(¯σμν)˙α˙β¯Q˙β.
We have to write down the anticommutation relation between the Grassman objects Qα
and¯Q˙βand now the work we did in appendix E really pays off. The supersymmetry algebra
is given by
{Qα,¯Q˙β}=2(σμ)α˙βPμ (1)
2One complex degree of freedom on mass shell and two complex degrees of freedom off mass shell. See the
superfield formalism below.
VIII.4. Supersymmetry | 463
We argue by the “what else can it be?” method. The right-hand side must carry the indices
αand˙βand we know that the only object that carries these indices is σμ. The Lorentz
index μhas to be contracted and the only vector around is Pμ. The factor of 2 fixes the
normalization of Q.
By the same kind of argument we must have {Qα,Qβ}=c1(σμν)αβJμν+c2δβ
α. Com-
muting with Pλwe see that the constant c1must vanish. Recalling that Qγ=εγβQβ,w e
have{Qα,Qγ}=c 2εγα; but since the left-hand side is symmetric in αandγwe have c2=0.
Thus, {Qα,Qβ}=0 and {¯Q˙α,¯Q˙β}=0 (see exercise VIII.4.2).
A basic theorem
An important physical fact follows immediately from (1). Contracting with (¯σν)˙βαwe
obtain
4Pν=(¯σν)˙βα{Qα,¯Q˙β} (2)
In particular the time component tells us about the Hamiltonian
4H=/summationdisplay
α{Qα,¯Q˙α}=/summationdisplay
α{Qα,Q†
α}=/summationdisplay
α(QαQ†
α+Q†
αQα) (3)
We obtain the important theorem that in a supersymmetric field theory any physical state
|S/angbracketrightmust have nonnegative energy:
/angbracketleftS|H|S/angbracketright=1
2/summationdisplay
α/summationdisplay
S/prime|/angbracketleftS/prime|Qα|S/angbracketright|2≥0 (4)
Superspace
Now that we have constructed the supersymmetric algebra let us keep in mind our goal of
constructing supersymmetric field theories. T o do that, we need to figure out and classifyhow fields transform under this supersymmetric algebra. We have to go through a lot offormalism, the necessity for which will become clear in due course.
Imagine that you are trying to invent the superspace formalism. Let us motivate it by
staring at the basic relation (1) {Q
α,¯Q˙β}=2(σμ)α˙βPμ. A supersymmetric transformation
Qfollowed by its conjugate ¯Q˙βgenerates a translation Pμ. Hmm, let’s see, Pμ≡i(∂/∂xμ)
generates translation in xμ, so perhaps Qα, being Grassmannian, would generate transla-
tion in some abstract Grassmannian coordinate θα? (Similarly, ¯Q˙βwould generate trans-
lation in ¯θ˙β.)
Salam and Strathdee invented the notion of a superspace with bosonic and fermionic
coordinates {xμ,θα,¯θ˙β}with the supersymmetry algebra represented by translations in
this space.
So let us try Qαand¯Q˙βbeing something like ∂/∂θαand∂/∂¯θ˙β, respectively. But then
{Qα,¯Q˙β}=0 and we don’t get (1). We have to keep playing around modifying Qαand¯Q˙β.
You may already see what we need. If we add a term such as θσμ∂μto¯Q˙β, then the ∂/∂θα
464 | VIII. Gravity and Beyond
inQαacting on θσμ∂μwill produce something like the right-hand side of (1). Similarly,
we will want to add a term such as ¯θσμ∂μtoQα. (Once again, the dotted and undotted
notation we worked hard to develop fixes what we must write, namely (σμ)α˙α¯θ˙α∂μso that
the indices match and obey the “southwest to northeast” rule.) Thus, we represent thesupercharges as
Qα=∂
∂θα−i(σμ)α˙α¯θ˙α∂μ (5)
and
¯Q˙β=−∂
∂¯θ˙β+iθβ(σμ)β˙β∂μ (6)
You see that (1) is now satisfied. Interestingly, when we translate in the fermionic direction
we have to translate a bit in the bosonic direction as well.
Superfield
A superfield /Phi1(xμ,θα,¯θ˙β), as the name suggests, is just a field living in superspace. An
infinitesimal supersymmetry transformation takes
/Phi1→/Phi1/prime=(1+iξαQα+i¯ξ˙α¯Q˙α)/Phi1 (7)
withξand¯ξtwo Grassmannian parameters.
It turns out that we can impose some condition on /Phi1and restrict this rather broad
definition a bit. After staring at (5) and (6) for a while, you may realize that there are twoother objects,
Dα=∂
∂θα+i(σμ)α˙α¯θ˙α∂μ
and
¯D˙β=−/bracketleftbigg∂
∂¯θ˙β+iθβ(σμ)β˙β∂μ/bracketrightbigg
that we can define, sort of the combinations orthogonal to Qαand¯Q˙β. Clearly, Dαand
¯D˙βanticommute with Qαand¯Q˙β. The significance of this fact is that if we impose the
condition ¯D˙β/Phi1=0 on the superfield /Phi1, then according to (7) its transform /Phi1/primealso satisfies
the condition.
A superfield /Phi1satisfying the condition ¯D˙β/Phi1=0 is known as a chiral superfield. The con-
dition is actually easy to implement:3Observe that if we define yμ≡(xμ+iθα(σμ)α˙α¯θ˙α)
(note we are adding two bosonic quantities here), then
¯D˙βyμ=−/bracketleftbigg∂
∂¯θ˙β+iθβ(σν)β˙β∂ν/bracketrightbigg
yμ=−/bracketleftBig
−iθα(σμ)α˙β+iθβ(σμ)β˙β/bracketrightBig
=0
Thus, a superfield /Phi1(y ,θ)that depends on yandθonly is a chiral superfield.
3This is analogous to the problem of constructing a function f( x ,y)satisfying the condition Lf=0 with
L≡[x(∂/∂y) −y(∂/∂x) ]. We define r≡(x2+y2)1
2and observe that Lr=0. Then any fthat only depends on
rsatisfies the desired condition.
VIII.4. Supersymmetry | 465
Let us expand /Phi1in powers of θholding yfixed. Remember that θcontains two compo-
nents (θ1,θ2). Thus, we can form an object with at most two powers of θ, namely θθ, which
you worked out in exercise E.3. Thus, as usual, power series in Grassmannian variablesterminate, and we have
/Phi1(y ,θ)=ϕ(y)+√
2θψ(y) +θθF(y)
withϕ(y) ,ψ(y) , andF(y) merely coefficients in the series at this stage. We can T aylor
expand once more around x:
/Phi1(y ,θ)=ϕ(x)+√
2θψ(x) +θθF(x)
+iθσμ¯θ∂μϕ(x)−1
2θσμ¯θθσν¯θ∂μ∂νϕ(x)+√
2θiθσμ¯θ∂μψ(x)(8)
We see that a chiral superfield /Phi1contains a Weyl fermion field ψ, and two complex scalar
fields ϕandF.
Finding a total divergence
Let’s do a bit of dimensional analysis for fun and profit. Given that Pμhas the dimension
of mass, which we write as [Pμ]=1 using the same notation as in chapter III.2, then (1),
(5), and (6) tell us that [ Q]=[¯Q]=1
2and [θ]=[¯θ]=−1
2. Given [ ϕ]=1, then (8) tells us
[ψ]=3
2, which we know already, and [ F]=2, which we didn’t know. In fact, we have never
met a Lorentz scalar field with mass dimension 2. How can we have a kinetic energy termforFinLwith dimension 4? We can’t. The term F
†Falready has dimension 4, and any
derivative is going to make the dimension even higher. Also, didn’t we say that with ϕand
ψwe balance the same number of bosonic and fermionic degrees of freedom?
The field F(x) definitely has something strange about him. What is he doing in our
theory?
Under an infinitesimal supersymmetric transformation the superfield changes by δ/Phi1=
i(ξQ+¯ξ¯Q)/Phi1 . Referring to (8), (5), and (6), you can work out how the component fields ϕ,
ψ, andFtransform (see exercise VIII.4.5). But we can go a long way invoking symmetry
and dimensional analysis. For example, δF is linear in ξor¯ξ, which by dimensional
analysis must multiply something with dimension [5
2] since [ F]=2 and [ ξ]=[¯ξ]=−1
2.
The only thing around with dimension [5
2]i s∂μψ, which carries an undotted index. Note it
can’t be ∂μ¯ψsince/Phi1does not contain ¯ψ. By Lorentz invariance we have to find something
carrying the index μ, and that can only be (σμ)α˙α. The dotted index on (σμ)α˙αcan only be
contracted with ¯ξ. So everything is fixed except for an overall constant:
δF∼∂μψα(σμ)α˙α¯ξ˙α(9)
Arguing along the same lines you can easily show that δϕ∼ξψ andδψ∼ξF+∂μϕσμ¯ξ.
The important point here is not the overall constant in (9) but that δF is a total diver-
gence.
Given any superfield /Phi1let us denote by [ /Phi1]Fthe coefficient of θθin an expansion of /Phi1
[as in (8)]. What we have learned is that under a supersymmetric transformation δ([/Phi1]F)
is a total divergence and thus/integraltext
d4x[/Phi1]Fis invariant under supersymmetry.
466 | VIII. Gravity and Beyond
Our next observation is that if ¯D˙β/Phi1=0, then ¯D˙β/Phi12=0 also. In other words, if /Phi1is a
chiral superfield, then so is /Phi12(and by extension, /Phi13,/Phi14, and so forth).
Supersymmetric action
What do we want to achieve anyway? We want to construct an action invariant under
supersymmetry.
Finally, after all this formalism we are ready. In fact, it is almost staring us in the face:/integraltext
d4x[1
2m/Phi12+1
3g/Phi13+...]Fis invariant under supersymmetry by virtue of the last two
paragraphs. Squaring (8) and extracting the coefficient of θθwe see by inspection that
[/Phi12]F=(2Fϕ−ψψ) . Similarly, [ /Phi13]F=3(Fϕ2−ϕψψ) . Now do exercise VIII.4.6.
Looks like we have generated a mass term for the Weyl fermion ψand its coupling to
the scalar field ϕ, but where are the kinetic energy terms, such as ¯ψ˙α(¯σμ)˙αα∂μψα?
Vector superfield
The kinetic energy terms contain ¯ψ˙α, which does not appear in /Phi1. T o get the conjugate
field¯ψ˙α, we obviously have to use /Phi1†, and so we are led to consider /Phi1†/Phi1. More formalism
here! We call a superfield V( x ,θ,¯θ)a vector superfield if V=V†. For example, /Phi1†/Phi1is a
vector superfield.
Imagine expanding /Phi1†/Phi1=ϕ†ϕ+... or any vector superfield Vin powers of θand¯θ.
The highest power is uniquely ¯θ¯θθθ since by the properties of Grassmannian variables the
only object we can form is ¯θ˙1¯θ˙2θ1θ2. Any object quadratic in θand quadratic in ¯θ, such as
(θσμ¯θ)(θσ μ¯θ), can be beaten down to ¯θ¯θθθ by using the kind of identities you discovered
in the exercises in appendix E. Let [ V]Ddenote the coefficient of ¯θ¯θθθ in the expansion
ofV.
Again, dimensional analysis can carry us a long way. If Vhas mass dimension n,
then [ V]Dhas mass dimension n+2 since θand¯θeach has mass dimension −1
2. Let
us study how [ V]Dchanges under an infinitesimal supersymmetry transformation δV=
i(ξQ+¯ξ¯Q)V . We use the same kind of argument as before: δ([V]D)is linear in ξor¯ξ,
which by dimensional analysis must multiply something with dimension n+5
2since [ ξ]=
[¯ξ]=−1
2. This can only be the derivative ∂of something with dimension n+3
2, namely
the coefficients of ¯θ¯θθand¯θθθ in the expansion of V. We conclude that δ([V]D)has to
have the form ∂μ(...), namely that δ([V]D)is a total divergence. This is the same type of
argument that allows us to conclude that δ([/Phi1]F)is a total divergence.
Thus, the action/integraltext
d4x[/Phi1†/Phi1]Dis invariant under supersymmetry.
Staring at (8), which I repeat for your convenience,
/Phi1(y ,θ)=ϕ(x)+√
2θψ(x) +θθF(x) (10)
+iθσμ¯θ∂μϕ(x)−1
2θσμ¯θθσν¯θ∂μ∂νϕ(x)+√
2θiθσμ¯θ∂μψ(x)
VIII.4. Supersymmetry | 467
we see that/integraltext
d4x[/Phi1†/Phi1]Dcontains/integraltext
d4xϕ†∂2ϕ(from multiplying the first term in /Phi1†with
the fifth term in /Phi1),/integraltext
d4x∂ϕ†∂ϕ(from multiplying the fourth term in /Phi1†with the fourth
term in /Phi1),/integraltext
d4x¯ψ¯σμ∂μψ(from multiplying the second term in /Phi1†with the sixth term
in/Phi1), and finally/integraltext
d4xF†F(from multiplying the third term in /Phi1†with the third term
in/Phi1). It is quite amusing how derivatives of fields arise in supersymmetric field theories:
Note that the action/integraltext
d4x[/Phi1†/Phi1]Ddoes not contain derivatives explicitly.
T o summarize, given a chiral superfield /Phi1we have constructed the supersymmetric
action
S=/integraldisplay
d4x{[/Phi1†/Phi1]D−([W(/Phi1)]F+h.c.)} (11)
Explicitly, with the choice W(/Phi1) =1
2m/Phi12+1
3g/Phi13, we have
S=/integraldisplay
d4x{∂ϕ†∂ϕ+i¯ψ¯σμ∂μψ+F†F−(mFϕ −1
2mψψ +gFϕ2−gϕψψ +h.c.)} (12)
An auxiliary field
From the very beginning the field Fseemed strange. Since [ F]=2 we anticipated that it
cannot have a kinetic energy term with mass dimension 4 and indeed it doesn’t. We see thatit is not a dynamical field that propagates—it is an auxiliary field (just like σin chapter III.5
andξ
μin chapter VI.3) and can be integrated out in the path integral/integraltext
DF†DFeiS. Indeed,
collect the terms that depend on FinS, namely
F†F−F(mϕ +gϕ2)−F†(mϕ†+gϕ†2)=|F−(mϕ+gϕ2)†|2−|mϕ+gϕ2|2
So, integrate over FandF†and get
S=/integraldisplay
d4x{∂ϕ†∂ϕ+i¯ψ¯σμ∂μψ−|mϕ+gϕ2|2+(1
2mψψ −gϕψψ +h.c.)} (13)
Note that the scalar potential V( ϕ†,ϕ)=|mϕ+gϕ2|2≥0 in accordance with (4) and
vanishes at its minimum, giving a zero cosmological constant. Note that we are no longerfree to add an arbitrary constant to V( ϕ
†,ϕ)as we could in a nonsupersymmetric field
theory.
As expected, supersymmetric field theories are much more restrictive than ordinary
field theories, and, duh, also much more symmetric. The formalism described here canbe extended to construct supersymmetric Y ang-Mills theory.
Another important generalization is to introduce, instead of one supercharge Q
α,Nsu-
percharges QI
α, with I=1,... ,N(exercise VIII.4.2). Since each charge QI
αtransforms
like the Sz=1
2component of a spin1
2operator, it takes a state with Sz=min a super-
multiplet to a state with Sz=m+1
2. Thus the integer Nis bounded from above. For
supersymmetric Y ang-Mills theory, the maximum number of supersymmetry generatorsisN=4 if we do not want to introduce fields with spin ≥1. Similarly, the most supersym-
metric supergravity theory we could construct (exercise VIII.4.3) has N=8.
468 | VIII. Gravity and Beyond
As mentioned in chapter VII.3, if any nontrivial 4-dimensional quantum field theory
turned out to be exactly soluble, the supersymmetric N=4 Y ang-Mills theory is probably
our best bet. In all likelihood, the first relativistic quantum field theory to be solved exactlywould be N=4 Y ang-Mills in the planar large Nlimit of chapter VII.4.
I hope that this brief introduction gave you a flavor of supersymmetry and will enable
you to go on to specialized treatises.
Exercises
VIII.4.1 Construct the Wess-Zumino Lagrangian by the trial and error approach.
VIII.4.2 In general there may be Nsupercharges QI
α, with I=1 ,..., N. Show that we can have {QI
α,QJ
β}=
εαβZIJ, where ZIJdenotes c-numbers known as central charges.
VIII.4.3 From the fact that we do not know how to write consistent quantum field theories with fields having
spin greater than 2 show that the Nin the previous exercise cannot exceed 8. Theories with N=8
supersymmetry are said to be maximally supersymmetric. Show that if we do not want to include gravity,Ncannot be greater than 4. Supersymmetric N=4 Y ang-Mills theory has many remarkable properties.
VIII.4.4 Show that ∂θ
α/∂θβ=εαβ.
VIII.4.5 Work out δϕ,δψ, andδFprecisely by computing δ/Phi1=i(ξαQα+¯ξ˙α¯Q˙α)/Phi1.
VIII.4.6 For any polynomial W(/Phi1) show that [W(/Phi1)]F=F[dW(ϕ)/dϕ ]+terms not involving F. Show that for
the theory (11) the potential energy is given by V( ϕ†,ϕ)=|∂W(ϕ)/∂ϕ |2.
VIII.4.7 Construct a field theory in which supersymmetry is spontaneously broken. [Hint: You need at least three
chiral superfields.]
VIII.4.8 If we can construct supersymmetric quantum field theory, surely we can construct supersymmetric
quantum mechanics. Indeed, consider Q1≡1
2[σ1P+σ2W(x) ] andQ2≡1
2[σ2P−σ1W(x) ], where the
momentum operator P=−i(d/dx) as usual. Define Q≡Q1+iQ2. Study the properties of the Hamil-
tonian Hdefined by {Q,Q†}=2H.
VIII.5A Glimpse of String Theory
as a 2-Dimensional Field Theory
Geometrical action for the bosonic string
In this chapter, I will try to give you a tiny glimpse into string theory. Needless to say, you
can get only the merest whiff of the subject here, but fortunately excellent texts do exist andI believe that this book has prepared you for them. My main purpose is to show you thatperhaps surprisingly the basic formulation of string theory is naturally phrased in termsof a 2-dimensional field theory.
In chapter I.11 I described a point particle tracing out a world line given by X
μ(τ) in
D−dimensional spacetime. Recall that the action is given geometrically by the length of
the world line
S=−m/integraldisplay
dτ/radicalBigg
dXμ
dτdXμ
dτ(1)
and remains unchanged under reparametrization τ→τ/prime(τ). Recall also that classically, S
is equivalent to
Simp=−1
2/integraldisplay
dτ/parenleftbigg1
γdXμ
dτdXμ
dτ+γm2/parenrightbigg
(2)
Now consider a string sweeping out a world sheet given by Xμ(τ,σ)inD-dimensional
spacetime, which we have already encountered in chapter IV .4 in connection with differ-
ential forms. In analogy with (1), Nambu and Goto proposed an action given geometricallyby the area of the world sheet
SNG=T/integraldisplay
dτdσ/radicalBig
det(∂aXμ∂bXμ) (3)
where ∂1Xμ≡∂Xμ/∂τ ,∂2Xμ≡∂Xμ/∂σ , and(∂aXμ∂bXμ)denotes the abelement of a 2 by
2 matrix. Here, as in (1), μranges over Dvalues: 0, 1, . . . , D−1. The constant T(≡1/2πα/prime
withα/primethe slope of the Regge trajectory in particle phenomenology) corresponds to the
string tension since stretching the string to enlarge the world sheet costs an extra amountof action proportional to T.
470 | VIII. Gravity and Beyond
In a precise parallel with the discussion for the point particle, it is preferable to avoid
the square root and instead use the action
S=1
2T/integraldisplay
dτdσγ1
2γab(∂aXμ∂bXμ) (4)
withγ=detγabin the path integral to quantize the string. We will now show that Sis
equivalent classically to SNG.
As in (2), we vary Swith respect to the auxiliary variable γab, which we then eliminate.
For a matrix M,δM−1=−M−1(δM)M−1andδdetM=δetr logM=etr logMtrM−1δM=
(detM)trM−1δM. Thus, δγab=−γacδγcdγdbandδγ=γγbaδγab. For ease of writing,
define hab≡∂aXμ∂bXμ. The variation of the integrand in (4) thus gives
δ[γ1
2γabhab]=γ1
2[1
2γdcδγcd(γabhab)−γacδγcdγdbhab]
Setting the coefficient of δγcdequal to 0 we obtain
hcd=1
2γcd(γabhab) (5)
where the indices on hare raised and lowered by the metric γ. Multiplying (5) by hdc(and
summing over repeated indices) we find γabhab=2 and thus γcd=hcd. Plugging this into
(4) we find that S=T/integraltext
dτdσ( deth)1
2. Thus, SandSNGare indeed equivalent classically.
The action (4), first discovered by Brink, Di Vecchia, and Howe and by Deser and Zumino,is known as the Polyakov action.
Note that (5) determines γ
abonly up to an arbitrary local rescaling known as a Weyl
transformation:
γab(τ,σ)→e2ω(τ ,σ)γab(τ,σ) (6)
Thus, the action (4) must be invariant under the Weyl transformation.
Staring at the string action (4), you will recognize that it is just the action for a quantum
field theory of Dmassless scalar fields Xμ(τ,σ)in 2-dimensional spacetime with coor-
dinates (τ,σ), albeit with some unusual signs. The index μplays the role of an internal
index, and Poincar ´e invariance in our original D-dimensional spacetime now appears as
an internal symmetry. Indeed, a good deal of string theory is devoted to the study of quan-tum field theories in 2-dimensional spacetime! It is amusing how quantum field theorymanages to stay on the stage.
T o this bosonic string theory we can add fermionic variables in such a way as to make
the action supersymmetric. The result, as you surely have heard, is superstring theory,thought by some to be the theory of everything.
1
1“T o understand macroscopic properties of matter based on understanding these microscopic laws is just
unrealistic. Even though the microscopic laws are, in a strict sense, controlling what happens at the larger scale,they are not the right way to understand that. And that is why this phrase, “theory of everything,” sounds sleazy.”—
J. Schwarz, one of the founders of string theory.
VIII.5. A Glimpse of String Theory | 471
This infinitesimal introduction to string theory is all I can give you here, but I hope that
this book has prepared you adequately to begin studying various specialized texts on stringtheory.
2
2For a brief but authoritative introduction, see E. Witten, “Reflections on the Fate of Spacetime,” Physics T oday ,
April 1996, p. 24.
This page intentionally left blank
Closing Words
As I confessed in the preface, I started out intending to write a concise introduction to
quantum field theory, but the book grew and grew. The subject is simply too rich. As Imentioned, after a period of almost being abandoned, quantum field theory came roaringback. T o quote my thesis advisor Sidney Coleman, the triumph of quantum field theorywas veritably “a victory parade” that made “the spectator gasp with awe and laugh withjoy.”
String theory is beautiful and marvellous, but until it is verified, quantum field theory
remains the true theory of everything. All of physics can now be said to be derivable fromfield theory. T o start with, quantum field theory contains quantum mechanics as a (0 +1)-
dimensional field theory, and to end (perhaps) with, string theory may be formulated as a(1+1)-dimensional field theory.
Quantum field theory can arguably be regarded as the pinnacle of human thought.
(Hush, you hear the distant howls of the mathematicians, English professors, philoso-phers, and perhaps even a few stuck-up musicologists?) It is a distillation of basic notionsfrom the very beginning of the physics: Newton’s realization that energy is the square ofmomentum appears in field theory as the two powers of spatial derivative. But yet—youknew that was coming, didn’t you, with field theory set up as the pinnacle et cetera?—but yet, field theory in its present form is in my opinion still incomplete and surely somebright young minds will see how to develop it further.
For one thing, field theory has not progressed much beyond the harmonic paradigm, as
I presaged in the first chapter. The discovery of the soliton and instanton opened up a newvista, showing in no uncertain terms that Feynman diagrams ain’t everything, contraryto what some field theorists thought. Duality offers one way of linking perturbative weak
coupling theory to strong coupling, but as yet practically nothing is known of the strongcoupling regime. When speaking of renormalization groups, we bravely speak of flowingto a strong coupling fixed point, but we merely have the boat ticket: We have little idea of
474 | Closing Words
what the destination looks like. Perhaps in the not too distant future, lattice field theorists
can extract the field configurations that dominate.
Another restriction is to two powers of the derivative, a restriction going back to Newton
as I remarked above. In modern applications of field theory to problems far beyond particlephysics, there is no reason at all to impose this restriction. For example, in studying visualperception, one encounters field theories much more involved than those we have studiedin this book. (See the appendix for a brief description.) These field theories are Euclideanin any case and the corresponding functional integral with higher derivatives certainlymakes sense: It is only in Minkowskian theories that we do not know how to handlehigher derivatives. Newton again—certainly economists consider the rate of change ofthe acceleration as well as acceleration. Another innovative application is the formulation
1
of a class of problems in nonequilibrium statistical mechanics as field theories. Typically,various objects wander around and react when they meet. This class of problems appearsin areas ranging from chemical reactions to population biology.
We can go far beyond the restriction on the number of derivatives in the Lagrangian.
Who said that we can only have integrands of the form “exponential of a spacetimeintegral”? Most modifications you can think of might immediately run afoul of some basic
principles (for example,/integraltext
Dϕe
−/integraltext
d4xL(ϕ)−[/integraltext
d4x˜L(ϕ)]2would violate locality), but surely
others might not. Another speculative thought I like to entertain goes along the followingline: Classical and quantum physics are formulated in terms of differential equationsand functional integrals, respectively. But how are differential equations contained in
integrals? The answer is that the integrals/integraltext
Dϕe
−(1//planckover2pi)/integraltext
d4xL(ϕ)contain a parameter /planckover2piso
that in the limit /planckover2pigoing to zero the evaluation of the integrals amounts to solving partial
differential equations. Can we go beyond quantum field theory by finding a mathematicaloperation that in the limit of some parameter
¯kgoing to zero reduces to doing the integral/integraltext
Dϕe−(1//planckover2pi)/integraltext
d4xL(ϕ)?
The arena of local field theory has always been restricted to the set of dreal numbers
xμ. The recent excitement over noncommutative field theory promises to take us beyond.
(I was tempted to discuss noncommutative field theory too, but then the nutshell wouldtruly burst.)
But perhaps the most unsatisfying feature of field theory is the present formulation of
gauge theories. Gauge “symmetry” does not relate two different physical states, but twodescriptions of the same physical state. We have this strange language full of redundancywe can’t live without. We start with unneeded baggage that we then gauge-fix away. We evenknow how to avoid this redundancy from the start but at the price of discretizing spacetime.This redundancy of description is particularly glaring in the manufactured gauge theoriesnow fashionable in condensed matter physics, in which the gauge symmetry is not thereto begin with. Also, surely the way we calculate in nonabelian gauge theories by cuttingthe Y ang-Mills action up into pieces and doing violence to gauge invariance will be held
1By M. Doi, L. Peliti, J. C. Cardy, and others. See for example J. C. Cardy, cond-mat/9607163, “Renormalisation
Group Approach to Reaction-Diffusion Problems,” in: J.-B. Zuber, ed., Mathematical Beauty of Physics , p. 113.
Closing Words | 475
up to ridicule a hundred years from now. I would not be surprised if a brilliant reader of
this book finds a more elegant formulation of what we now call gauge theories.
Look at the development of the very first field theory, namely Maxwell’s theory of
electromagnetism. By the end of the nineteenth century it had been thoroughly studied andthe overwhelming consensus was that at least the mathematical structure was completelyunderstood. Yet the big news of the early twentieth century was that the theory, surprisesurprise, contains two hidden symmetries, Lorentz invariance and gauge invariance: twosymmetries that, as we now know, literally hold the key to the secrets of the universe. Mightnot our present day theory also contain some unknown hidden symmetries, symmetrieseven more lovely than Lorentz and gauge invariance? I think that most physicists wouldsay that the nineteenth-century greats missed these two crucial symmetries because oftheir lousy notation
2and tendency to use equations of motion instead of the action. Some
of these same people would doubt that we could significantly improve our notation andformalism, but the dotted-undotted notation looks clunky to me and I have a naggingfeeling
3that a more powerful formalism will one day replace the path integral formalism.
Since the point of good pedagogy is to make things look easy, students sometimes do
not fully appreciate that symmetries do not literally leap out at you. If someone had writtena supersymmetric Y ang-Mills theory in the mid-1950s, it would certainly have been a longtime before people realized that it contained a hidden symmetry. So it is entirely possiblethat an insightful reader could find a hitherto unknown symmetry hidden in our well-studied field theories.
It is not just a matter of clearer notation and formalism that caused the nineteenth-
century greats to miss two important symmetries; it is also that they did not possessthe mind set for symmetry. The old paradigm “experiments →action →symmetry” had
to be replaced
4in fundamental physics by the new paradigm “symmetry →action →
experiments,” the new paradigm being typified by grand unified theory and later by stringtheory. Surely, some future physicists will remark archly that we of the early twenty-firstcentury did not possess the right mind set.
In physics textbooks, many subjects have a finished completed feel to them, but not
quantum field theory. Some people say to me, what else is there to say about field theory?I would like to remind those people that a large portion of the material in this book wasunknown 30 years ago. Of course, while I feel that further developments are possible, Ihave no idea what—otherwise I would have published it—so I can’t tell you what. But let memention two recent developments that I find extremely intriguing. (1) Some field theoriesmay be dual to string theories. (2) In dimensional deconstruction a d-dimensional field
theory may look (d+1)-dimensional in some range of the energy scale: the field theory
can literally generate a spatial dimension. These developments suggest that quantum field
2It is said, and I agree, that one of Einstein’s great contributions is the repeated indices summation convention.
T ry to read Maxwell’s treatises and you will appreciate the importance of good notation.
3I once asked Feynman how he would solve the finite square well using the path integral.
4A. Zee, Fearful Symmetry, chap. 6.
476 | Closing Words
theories contain considerable hidden structures waiting to be uncovered. Perhaps another
golden age is in store for quantum field theory.
So boys and girls, the parade is over, and now it’s up to you to get another parade going.5
Appendix
An image presented to the visual system can be described as a 2-dimensional Euclidean field ϕ0(x), with ϕ0
representing the gray scale from black ( ϕ0=− ∞ )to white ( ϕ0=+ ∞ ). [You can see that color might be included
by going to a field /vectorϕtransforming under some internal SO( 2)group for example.] The image actually perceived,
ϕ(x) , is the actual image ϕ0(x) distorted to ϕ0[y(x) ] plus some noise η(x) . Distortion is described by a map
x→y(x) of the 2-dimensional Euclidean plane. Your brain’s task is to decide whether the actual image is ϕ0(x)
or some other ϕ1(x). Your ability to discriminate between images depends on the functional integral
Z=/integraldisplay
Dy(x)/integraldisplay
Dη(x)e−W [y(x) ]−(1/2C)/integraltext
d2xη(x)2δ{ϕ0[y(x) ]+η(x)−ϕ(x)} (1)
=/integraldisplay
Dy(x)e−W [y(x) ]−(1/2C)/integraltext
d2x{ϕ(x)−ϕ 0[y(x) ]}2
where for simplicity I have taken the noise, measured by the parameter C, to be Gaussian and white. The
weighting function W[y(x) ] is presumably hard wired by evolution into our visual system, telling us that certain
distortions (translations, rotations, and dilations) are much more likely than others. Writing y(x)=x+A(x) we
note that Zdefines a field theory of the 2-component field Ai(x), which can always be written as Ai=∂iη+εij∂jχ.
Note that the field Ai(x) appears “inside” an “external” field ϕ0. From symmetry considerations we might argue
that
W=−/integraldisplay
d2x/parenleftbigg1
g2η∂6η+1
f2χ∂6χ/parenrightbigg
with two coupling constants fandg. I can give here only the briefest of sketches and refer the interested reader
to the literature.6Clearly, one can think of other examples. This particular example serves only to show that there
are many more field theories than those described in standard texts.
5As the Beatles said, quantum fields forever!
6W . Bialek and A. Zee, Statistical mechanics and invariant perception, Phys. Rev. Lett. 58: 741, 1987; Under-
standing the efficiency of human perception, Phys. Rev. Lett. 61: 1512, 1988.
Part N
While quantum field theory was discovered and developed in the twentieth century, I will
introduce in this part, added in the second edition of this text, some topics that have beenworked out in the twenty-first century. At the rate these topics are rapidly evolving, I maybe quite foolish to include them here. But I am taking the plunge as I think that I wouldserve my readers better by letting them have a taste of the twenty-first century ratherthan expanding on the twentieth. More likely than not, by the time this second editionis published, there will be better ways of treating the material contained here. You shouldread part N in this spirit and regard what is given here as an entry key to a fast growing
1
research literature.
1Indeed, by the time the manuscript was copyedited (April 2009) it had been discovered that the amplitudes
discussed in chapters N.2–4 could be written even more simply using a twistor and dual twistor formalism. See
p. 494 and N. Arkani-Hamed, P. Cachazo, C. Cheung, and J. Kaplan, arXiv:09032110.
This page intentionally left blank
N.1 Gravitational Waves and Effective Field Theory
An unfinished symphony
One astounding prediction of Einstein gravity is the existence of ripples crisscrossing the
fabric of spacetime, what one writer refers to as Einstein’s unfinished symphony.1Massive
detectors have been built, with more to come, in a human “curious George” effort to tunein to the song of the cosmos.
Consider a black hole of size r
S=2Gm (its Schwarzschild radius—see chapter I.10, with
mits mass) a distance rOfrom another black hole, moving with velocity v. As the black
holes spiral into each other they emit gravitational waves, with a characteristic wavelengthdetermined by the orbital period λ=2πr
O/v. Thus the physics contains three distance
scales: rS,rO, andλ. We will stay within the simple post Newtonian regime rS/lessmuchrO/lessmuchλ.
T oward the end, as rO/similarequalrSandv/similarequal1, relativistic effects rear their nasty heads, from which
we will prudently stay away.
In the Closing Words to the first edition of this text, I mention that one intriguing
development over the last few decades has been the use of effective field theory to describesituations involving more than one energy scale (or equivalently, length and time scales).The physics at the high energy scale Mis then represented in the low energy effective
Lagrangian by higher dimensional terms, suppressed by powers of Mbut constrained by
the symmetries we know. Examples abound in this book, from the quantum Hall effectto surface growth to proton decay. The latter provides a classic example: while we professignorance of the physics responsible for proton decay, we can nevertheless make usefulpredictions by adding 4-fermion interactions invariant under the low energy gauge groupSU( 3)⊗SU( 2)⊗U(1), as shown in chapter VIII.3.
An interesting recent development is an elegant description of the emission of gravi-
tational waves by inspiraling black holes using effective field theory. Here, in contrast toproton decay, we actually know the short distance physics involved. Effective field theory
1M. Bartusiak, Einstein ’s Unfinished Symphony.
480 | Part N
nevertheless offers an efficient and sensible way to organize and compartmentalize physics
on the various distance scales. We will merely touch upon one aspect of this approach.
Finite size objects in general relativity
SincerS/lessmuchrO, the leading approximation would be to treat the black hole as a point particle
using the action (I.11.12) Spp=−m/integraltext
dτ=−m/integraltext/radicalbiggμνdXμdXν=−m/integraltext
dτ/radicalBig
gμν˙Xμ˙Xν,
where ˙Xμ=dXμ/dτ . Let us now include the corrections due to the finite size of the black
hole. As you will see, the following discussion actually applies not only to black holes butto any finite sized object, including you.
In the spirit of effective field theory, we add to S
pphigher dimensional terms to be formed
out of the point particle degree of freedom ˙Xμand the ambient gμνthe particle moves in,
subject to local coordinate invariance, of course. The invariant tensors we can form out ofg
μνare, to leading order, the scalar curvature R, the Ricci curvature Rμν, and the Riemann
curvature tensor Rμλνρ . You might start with the scalar curvature and the Ricci curvature,
and add to Sppthe terms Sdrop=/integraltext
dτ(cSR(X) +cRRμν(X)˙Xμ˙Xν). The curvatures R(X)
andRμν(X) are evaluated on the worldline Xμ(τ) of the particle of course.
Einstein’s equation of motion Rμν−1
2gμνR=0 implies Rμν(X)=0 and thus also
R(X) =0. Following the discussion in chapter VIII.3, we now show that, as we might
intuitively feel, we are allowed to drop Sdrop. For our problem we have the total action
S=SEH+Spp+Sdrop with the Einstein-Hilbert action (VIII.1.1) SEH=/integraltext
d4x√−gM2
PR.
Under a field redefinition gμν→gμν+δgμν,
δSEH=/integraldisplay
d4x√−gM2
P(Rμν−1
2gμνR)δg μν (1)
Note that was how we would have derived the equation of motion for gravity, by varying gμν.
Here we are not making an arbitrary variation, but rather our goal is to choose a specificδg
μνso that the resulting δSEHnegates Sdrop. Since Sdrop consists of an integral over the
worldline of the particle while δSEHis given by an integral over spacetime, we need a delta
function in δgμνto switch from one kind of integral to another. The choice
δgμν(x)=1/radicalbig
−g(x)M2
P/integraldisplay
dτδ4(x−X(τ)) [agμν(X)+b˙Xμ˙Xν] (2)
gives
δS=/integraldisplay
dτ(−aR(X) +b/bracketleftbigg
Rμν(X)−1
2gμνR(X)/bracketrightbigg
˙Xμ˙Xν) (3)
So, with some appropriate values of aandb, we can indeed cancel off2Sdrop, thus
vindicating our intuition that the particle does not feel the Ricci and scalar curvaturesfor the obvious reason that they vanish.
2A technicality: field redefinition also induces a contact interaction of the form/integraltext
dτ( 1/√−g)δ4(X1(τ)−
X2(τ)) between the two massive objects. Going from field theory to the point particle description represents a
conceptual step backward, so that we should expect delta function effects at the location of the point particles.
N.1. Gravitational Waves | 481
How about terms we can construct out of the Riemann curvature tensor Rμλνρ ? For your
convenience, I list its symmetry properties3here:Rτρμν=−Rτρνμ=−Rρτμν ,Rτρμν=
Rμντρ , andRτρμν+Rτμνρ+Rτνρμ=0. Thus, due to the antisymmetry, we are not able to
contract all four indices of Rμλνρ with˙Xμ. We can contract at most two indices, to form
the two objects Eμν(X)≡Rμλνρ(X)˙Xλ˙XρandBμν(X)≡˜Rμλνρ(X)˙Xλ˙Xρ, where
˜Rμλνρ(x)≡1
2√−gεμλσηRση
νρ(x)
denotes the dual of the curvature tensor. These are 2-indexed tensors and we need to square
them to form scalars to put into the action. Hence, to next order the particle action becomes
Sp=/integraldisplay
dτ(−m+cEEμνEμν+cBBμνBμν+...) (4)
Note that the unknown constants cEandcBhave dimension of inverse mass cubed.
Before we explore the physical content of this effective action, let us understand the
meaning of EandBby retreating to the more familiar case of a point particle mov-
ing in an electromagnetic field Fμν(in flat space). Form Eμ≡Fμν˙XνandBμ≡˜Fμν˙Xν,
where ˜Fμν(x)≡1
2εμνσηFση. Going to the rest frame of the particle, where ˙X0=1 and
˙Xi=0, we see that, as the notation suggests, this is just the familiar decomposition of
electromagnetism into electricity and magnetism. Similarly, EμνandBμνrepresent the
decomposition of curvature into its “electric” and “magnetic” components.
In chapter I.11 we varied the first term in (4) to obtain the standard geodesic equation
that is at the heart of Einstein’s theory. Here we obtain
d2Xρ
dτ2+/Gamma1ρ
μν(X(τ))dXμ
dτdXν
dτ=fρ(X(τ)) (5)
where fμ(X(τ)) comes from varying the EandBterms in (4). A finite sized body
experiences a tidal force fμdue to the varying gravitational force acting on it. It no longer
follows a geodesic.
The fact that we had to square the electric and magnetic components of the curvature
to form the effective action (4) means that the effects of these correction terms are highlysuppressed. Since Riemann curvature contains two derivatives, the correction terms in-volve four derivatives. T o estimate the magnitude of c
EandcBwe exploit a rather cute
argument as follows.
Consider the scattering of a graviton off this point particle (which, remember, is a
black hole in the problem we are studying) generated by the couplings in (4): iM∼
...+icE, Bω4/M2
P+... where ωdenotes the energy of the graviton. The powers of ω
follows from the four derivatives just mentioned. (If you don’t understand the powersofM
Pyou need to read chapter VIII.1 again.) Here cE, B denotes the two unknown
couplings cE∼cBgenerically. Imagine calculating the total scattering cross section for
a graviton on a black hole. Squaring the amplitude Metc., we would end up with σ(ω)∼
...+c2
E, Bω8/M4
P+....
3S. Weinberg, Gravitation and Cosmology, p. 141.
482 | Part N
This treatment of the black hole as a point particle is only valid for ωrS/lessmuch1 of course. The
(...)iniMrepresents diagrams we have not included, for example, the one originating
from the first term in (4) (namely the term responsible for keeping us down to earth!). Anice feature of the argument I am about to give is that we don’t even need to know whatthe terms in (...)are.
On the other hand, we argue that by dimensional analysis the cross section must have
the form σ(ω)=r
2
Sf( ωr S)since the only length scale in the Schwarzschild metric is rS.
Expanding the unknown function f( ωr S)in powers of its argument we have σ(ω)=
...+αω8r10
S+...withαsome constant. (A technical aside: the massless graviton could
produce infrared factors like log ωrS, which we ignore for our purposes.)
Requiring that the two expressions agree, we obtain cE, B∼M2
Pr5
S. Indeed, as expected,
the couplings cE, Bare highly suppressed as rS→0.
Exercise
N.1.1 Using considerations similar to those in the text, show that the scattering cross section for a photon of
frequency ωon an atom or a molecule vanishes like ω4asω→0, a result which, as mentioned in chapter
VIII.3, underlies the well-known explanation of why the sky is blue.
N.2 Gluon Scattering in Pure Yang-Mills Theory
Boil and toil with Feynman diagrams
You might think that after some 50 years, there could not possibly be any novelty in
calculating Feynman amplitudes. But you would be wrong. Over the last dozen years or so,and largely since the first edition of this text, a group of intrepid searchers have found someamazingly powerful methods of tackling Feynman diagrams. As I said at the beginning ofchapters VIII.4 and 5, I can only give you an introduction to this subject, telling you justenough for you to explore this fast-growing literature.
T o best appreciate this new development, you should do a little calculation before reading
further. Consider pure Y ang-Mills theory, by consensus the nicest field theory we have,simple to write down and perfumed with symmetries. Not even any fermions around tomess things up. Call the gauge bosons gluons for convenience. Now calculate 5-gluonscattering at tree level as shown in figure N.2.1. No loops, just trees. The Feynman rulesare given in chapter VII.1 and also in appendix C.
You really must calculate before reading on. I will wait for you. You think to yourself,
this is easy, just a bunch of tree diagrams. In fact, to make it easier, put all the externalgluons on-shell, that is, set p
2
i=0,i=1, 2, ... ,5 .
This calculation is not merely an idle exercise, but is in fact phenomenologically im-
portant. At an accelerator such as the soon-to-be-operational Large Hadron Collider, two
/H11001 /H11001 /H11001...
Figure N.2.1
484 | Part N
Result of a brute force calculation (actually only a small part of it):
k1 /H11080 k4/H92552 /H11080 k1/H92551 /H11080 /H9255 3/H92554 /H11080 /H9255 5
Figure N.2.2
protons are smashed together at high energies. Two gluons, one from each proton, collide
and produce three gluons, which then materialize into three jets of hadrons. Because ofasymptotic freedom, at high energies the effective coupling gbecomes small enough for
perturbative field theory to be relevant, and the tree amplitude you are busily calculatingprovides a key ingredient for the phenomenological models used to study the experimentalmeasurements.
Time’s up! A small part of the answer is shown in figure N.2.2, taken from a lecture by
Zvi Bern.
1You really should take a look in order to appreciate, to be grateful for even, the
formalism to be explained in this chapter. You know that the amplitude is linear in eachof the five polarization vectors /epsilon1
i. The 3-gluon vertex (VII.1.11) is linear in momentum
and there are three of them in a typical diagram. Thus a typical term in the numerator
of the Feynman amplitude would be, as shown in the figure, p1.p4/epsilon12.p1/epsilon11./epsilon13/epsilon14./epsilon15.A
rough estimate shows that there are almost 10,000 such terms. That’s why, in spite of myadmonition, you didn’t finish the calculation before reading ahead. Incidentally, you couldsee that even 4-gluon scattering at tree level, though doable by hand, is rather involved.
1Z. Bern, “Magic T ricks for Scattering Amplitudes,” http://online.itp.ucsb.edu/online/colloq/bern1/pdf/
Bern1.pdf.
N.2. Gluon Scattering | 485
New technology for Feynman diagrams
In practice, phenomenologists studying jet production have developed elaborate computer
codes based on numerical recursion and these prove to be quite efficient. In this introduc-tory text, however, we are not after numerical efficiency but a deeper understanding of thestructure of multi-gluon amplitudes. I have set you up so that, surely, after your abortiveattempt to calculate the 5-gluon amplitude you now fully appreciate the need for new waysof approaching Feynman diagrams. I will now explain some of the novel methods peoplehave invented over the last 15 years or so.
A relatively simple first step is to strip the color off the amplitude. Evidently, it is much
better to use the matrix notation of (IV .5.16) than the index notation of (IV .5.17). Insteadof the structure constants f
abcand their products in the “indexed” Feynman rules in
(VII.1.11–13) we have colored Feynman rules (read appendix 1 now) with objects liketr(T
a[Tb,Tc])and tr ([Ta,Tb][Tc,Td])for the cubic and quartic coupling vertex, respec-
tively, where Tadenotes the matrix representing the suitably normalized generators of the
gauge group. Denote the color matrices carried by the external gluons by Ta1,Ta2,...Tan
(withn=5 in the example you failed to do).
In calculating a multi-gluon scattering amplitude in tree approximation, you would find
each term multiplied by the product of a bunch of color traces, such as tr (TeA)tr(TeB)
where AandBdenote products of T’s. Here the index eis carried by a virtual gluon and
hence is summed over. We now use the group theoretic identity (with esummed over) for
the gauge group SU(N) [recall (IV .5.19)]
(Te)i
j(Te)kl=1
2/parenleftbigg
δi
lδk
j−1
Nδi
jδk
l/parenrightbigg
(1)
(Here e=1,...,N2−1 and the indices i,j,k,l=1,...,N, of course.) The second term
takes care of the traceless condition tr Te=0. However, we can drop it, since if we extend
the gauge group to U(N) , that extra gluon does not couple to the other gluons anyway. Thus
tr(TeA)tr(TeB)=1
2tr(AB). Repeating this procedure, we reduce the product of traces to
a single trace of nTa’s multiplied together in some specific order.
Indeed, the astute reader will have noted that had we used the double-line formalism of
figure IV .5.2, this entire discussion would not even have been necessary. As we also sawin chapter VII.4, the double-line formalism does offer many advantages.
The other simplifying step is to specify the helicity of the gluon instead of writing
the amplitude in terms of polarization vectors. You recall, from way back when, that amassless spin 1 particle moving along the third-direction k=ω(1, 0, 0, 1 )can have helicity
h=+ , corresponding to the polarization vector /epsilon1=1/(√
2)(0, 1, i,0), or helicity h=− ,
corresponding to the polarization vector /epsilon1=(1/√
2)(0, 1, −i,0). We specify the external
gluons by momentum, helicity, and color: (p1,h1,a1,p2,h2,a2,... ,pn,hn,an).
Thus we can write the n-gluon amplitude as
M=i/summationdisplay
permutationstr(Ta1Ta2...Tan)A(1, 2, ... ,n) (2)
486 | Part N
Following the literature, we have compressed the notation further and denote {pi,hi}byi.
The sum is over all possible permutations of the ngluons. We can now focus on the “color
stripped” amplitude A(1, 2, ... ,n).
First, a triviality. It is convenient to treat the gluons as all outgoing (or if you prefer, as
all incoming), so that/summationtext
ipi=0 and the time component of some of the momenta can be
negative. We can then obtain the physically desired amplitude by crossing. Keep in mindthat under crossing p→−pand/epsilon1→/epsilon1
∗, that is, the helicity flips.
The spinor helicity formalism
Now we are ready to return to the expression in figure N.2.2. The technical term for this
expression is an “unholy mess.” It turns out that the key to unraveling this hopeless morasscan be found in exercise II.3.1: that the Lorentz vector sits in the representation (
1
2,1
2)and
thus can be constructed as a product of two spinors, one from the representation (1
2,0),
the other from (0,1
2). You did the exercise, didn’t you? So you know how to write, for
example, the momentum vector pμas a product of two spinors. T o go on, you should also
read appendices B and E.
I am now ready to explain the spinor helicity formalism designed to exploit this peculiar
property of the Lorentz vector. Or, to say it a bit more mysteriously, I am going to showyou how to take the square root of the momentum.
Now you appreciate the power of the undotted-dotted notation introduced in appendix E.
The undotted index goes with (
1
2,0), and the dotted with (0,1
2). We are looking for an object
transforming like (1
2,0)⊗(0,1
2)to represent a vector. The problem can then be stated as
follows: instead of writing momentum as pμ, we want to write it as pα˙α, an object carrying
an undotted and a dotted index, namelya2b y2matrix in cruder language.
We merely have to flip through appendix E and look for an object carrying the desired
indices. There it is, (σμ)α˙α, and indeed, its μindex is begging to be contracted with pμ.
Thus, with no further work, we can write [since σμ=(I,/vectorσ)]
pα˙α≡pμ(σμ)α˙α=(p0I−piσi)α˙α=/parenleftBigg(p0−p3)−(p1−ip2)
−(p1+ip2)( p0+p3)/parenrightBigg
α˙α(3)
We have succeeded in writing the momentum a sa2b y2matrix. You may recognize this
as nothing but the matrix XM(with some trivial change in notation) used in appendix B
to construct the covering of SO( 3, 1)bySL( 2,C).
Given two vectors pandq, their scalar product is given by
p.q=εαβε˙α˙βpα˙αqβ˙β (4)
which you can check explicitly, writing the right-hand side as a trace and once again using
σ2σT
i=−σiσ2, as we did in appendix E. For q=p, this reduces to p.p=εαβε˙α˙βpα˙αpβ˙β=
detp; here we recognized a definition of the determinant. [Of course, you could also eval-
uate the determinant of (3) by inspection, or recall that this was also used in appendix B.]
N.2. Gluon Scattering | 487
Clearly, there is an unavoidable notational overload: the single letter pdenotes both the
vector and the matrix, but you should be able to tell from the context which object is beingreferred to.
Here we are going to apply this formalism to massless gluons with lightlike momenta.
Things simplify considerably: for plightlike, det p=0 and thus the matrix pgenerically
has one 0 eigenvalue. (In fancy talk, the matrix has rank 1 rather than 2.) From elementarylinear algebra we recall thata2b y2matrix mof rank 1 can always be written as m
ij=viwj,
withvandwtwo 2-component vectors, (obviously since the vector orthogonal to wprovides
the 0 eigenvector.) Thus, for a lightlike vector, we can write
pα˙α=λα˜λ˙α (5)
in terms of two 2-component spinors λand˜λ.
For physical momentum, the components pμare real, of course. I invite you to verify,
however, that everything we just did from (3) to (5) goes through even if pμare complex.
It turns out that in the next chapters we will find it convenient to consider complexmomentum.
Upon first exposure, the formalism appears quite opaque, but actually, like a lot of
formalisms, it is fairly simple or perhaps even trivial. If you are confused at any point inthe following exposition, just work things out explicitly. For example, consider a physicalmomentum with p
0=E> 0. With no loss of generality, you can choose /vectorpto point along
the third direction, so that (with a trivial abuse of notation p=|/vectorp|)
p=/parenleftBiggE−p 0
0 E+p/parenrightBigg
which for plightlike collapses to the rank 1 matrix
p=2E/parenleftBigg00
01/parenrightBigg
=2E/parenleftBigg0
1/parenrightBigg
(01)
Thus, in this case, λand˜λare both equal to
√
2E/parenleftBigg0
1/parenrightBigg
numerically. (T o make sure you get it, work this out for /vectorppointing in some other direction.)
You can think of the Pauli spinors λand˜λas the “square root” of the Lorentz vector pμ.
Note how the group theory discussion in chapter II.3 foreordained this rather nontrivialpossibility. After all, there we saw how a Lorentz vector can be constructed out of two Diracspinors uandu
/prime.
Interestingly, in discussing ferromagnets and antiferromagnets in chapter VI.5, we used
a poor man’s version of (3), namely /vectorn=z†/vectorσz.
You learned in school that the ordinary square root has a sign ambiguity. Analogously,
in (5)pdoes not determine λand˜λuniquely. We can always rescale λ→uλand˜λ→1
u˜λ
for any complex number u. (You might have wondered what fixed the overall constant in
λand˜λin the simple example above: I made an arbitrary choice.)
488 | Part N
For real momentum, the matrix pα˙α=pμ(σμ)α˙αis hermitean, which implies that ˜λ=λ∗
is the complex conjugate of λ. The spinor ˜λis not independent of λ, and so the rescaling
parameter uis restricted to be a phase factor eiγ. [Also, recall from appendix B how XM
transforms under SL( 2,C) and you will see that it is all consistent.] In this case, the
condition that phas rank 1 allows for two solutions: pα˙α=±λα˜λ˙α, with the two possible
signs corresponding to whether p0>0 or not.
A side remark at this point: We will see that it is useful to consider the group SO( 2, 2)
instead of the Lorentz group SO( 3, 1). Thus, as the discussion in appendix B indicates,
you can also take the square root of an SO( 2, 2)vector and write pα˙α=λα˜λ˙α, but with λ
and˜λtwo independent real spinors, as is consistent with the local isomorphism between
SO( 2, 2)andSL( 2,R)⊗SL( 2,R). The rescaling mentioned above is now restricted to u
being a real number.
It is instructive to count the number of real degrees of freedom for these different
cases. A complex lightlike momentum depends on 4 ×2−2=6 real numbers, since the
condition p2now amounts to two real conditions, while λand˜λeach contains 2 complex
numbers, but with rescaling we are left with 2 ×2−1=3 complex numbers, that is, 6
real numbers. A real lightlike momentum depends on 4 −1=3 real numbers, but now
˜λis tied to λcontaining 2 complex numbers, which get reduced to 3 real numbers after
rescaling by a phase factor. For a (real) lightlike vector transforming under SO( 2, 2),w e
have 2 real spinors, which after rescaling contains 3 real numbers. So it all works out, ofcourse.
I mention all this here for future use. It should be evident to you, for the rest of
this chapter, which statements hold for complex momenta and which hold only for realmomenta. At the end of the day, when we arrive at a physical quantity, such as theamplitude, we will of course set the momenta contained therein to be real.
For two lightlike vectors pandq, write p
α˙α=λα˜λ˙αandqα˙α=μα˜μ˙α, then we have
p.q=(εαβλαμβ)(ε˙α˙β˜λ˙α˜μ˙β)≡/angbracketleftλ,μ/angbracketright[˜λ,˜μ] (6)
Here we have defined the two Lorentz invariants
/angbracketleftλ,μ/angbracketright≡εαβλαμβ=− /angbracketleftμ,λ/angbracketright (7)
and
[˜λ,˜μ]≡ε˙α˙β˜λ˙α˜μ˙β=− [˜μ,˜λ] (8)
(treating the spinors as c-number objects.) Note in passing that with our convention,
λ1=λ2andλ2=−λ1, and so /angbracketleftλ,μ/angbracketright=− λ1μ2+λ2μ1=−εαβλαμβ.
We have already verified in (E.13) that /angbracketleftλ,μ/angbracketrightis invariant, but for the sake of total
pedagogical clarity let us check it once more, this time using infinitesimal transformations.Write (E.4) more compactly as δλ
α=σβ
αλβ, where σdenotes some linear combination
of Pauli matrices. Noting that /angbracketleftλ,μ/angbracketrightis nothing but λσ2μup to some irrelevant overall
constant, we have indeed δ(λσ2μ)=(λσTσ2μ+λσ2σμ)=0.
A notational remark: the twiddles in [ ˜λ,˜μ] are redundant. The square bracket is defined
only for spinors transforming like (0,1
2). Henceforth, we will write [ λ,μ]≡ε˙α˙β˜λ˙α˜μ˙β.
N.2. Gluon Scattering | 489
For real physical momenta, ˜λ=λ∗so that /angbracketleftλ,μ/angbracketright=[ λ,μ]∗. Then p.q=/angbracketleftλ,μ/angbracketright[λ,μ]
implies that /angbracketleftλ,μ/angbracketright=√p.qeiφand [λ,μ]=√p.qe−iφ, with some phase factor eiφ.W e
thus conclude that the two spinorial products may be regarded as the (two) square rootsof the Lorentz dot product p.qup to a phase factor.
You could now raise an interesting question: how do we write the polarization vectors
/epsilon1(p) of a massless gluon?
The requirement that /epsilon1(p) .p=0 can be satisfied, according to (4), by setting /epsilon1
α˙α=
d−1λα˜μ˙α, for an arbitrary ˜μ˙αand with the factor ddetermined as follows. We require that,
for an arbitrary complex number w, scaling ˜μ→w˜μdoes not change /epsilon1(since ˜μis arbitrary
after all). Thus dhas to be linear in ˜μ. The further requirement that dbe Lorentz invariant
implies, as we just learned, that d=[x,μ], where ˜xis some (0,1
2)spinor. The only spinor
available is ˜λand hence we obtain
/epsilon1−
α˙α=λα˜μ˙α
[λ,μ](9)
By convention, we will call this polarization negative helicity.
The arbitrary choice of ˜μ˙αrepresents the freedom inherent in a gauge theory. Indeed,
we see that gauge transformation corresponds to the spinorial shift ˜μ→˜μ+y˜λ(for some
arbitrary number y) under which /epsilon1α˙α→/epsilon1α˙α+yλα˜λ˙α, which translates into the usual shift
of/epsilon1by some multiple of p.
The positive helicity polarization is given by the other possible choice
/epsilon1+
α˙α=μα˜λ˙α
/angbracketleftμ,λ/angbracketright(10)
Check that it works. Gauge transformation now corresponds to the shift μ→μ+yλ. Note
that the polarization vectors are normalized as /epsilon1+./epsilon1−=/angbracketleftμλ/angbracketright[μλ]/(/angbracketleftμλ/angbracketright[ μλ])=1.
Taming the unholy mess
Consider the tree-level scattering amplitude with n≥4 outgoing massless gluons. (In this
and the next sections, we can take all momenta to be real.) The color-stripped amplitudeis then characterized by a string of helicities (h
1,...,h n). T ake for example the amplitude
with(+++ ...++). Upon crossing, it describes two gluons, each with helicity −, going
inton−2 gluons all with helicity +. Both incoming gluons flip their helicity and thus
this amplitude is said to be maximal helicity violating. Your intuition may tell you thatthis amplitude ought to be suppressed, since highly energetic massless particles tend tomaintain their helicities. If you try to verify this using traditional Feynman diagrams, youwould once again encounter a big mess.
The spinor helicity formalism rides to the rescue. Consider the amplitude A(h
1,...,hn).
For each of the ngluons, we have piα˙α=λiα˜λi˙α, and an arbitrary spinor that we are
free to choose (subject to some conditions), namely either μiαor˜μi˙α, depending on
whether the corresponding helicity is +or−, respectively. There are quite a few indices,
but fortunately, in computing amplitudes, we encounter only Lorentz invariants, such as
490 | Part N
/epsilon1i./epsilon1j=εαβε˙α˙β/epsilon1iα˙α/epsilon1jβ˙β(be sure to distinguish between the two varieties of epsilon here!),
and thus the spinor indices will be contracted over and disappear. In particular, we have(omitting the comma in the angled and square brackets)
/epsilon1+
i./epsilon1+
j=/angbracketleftμiμj/angbracketright[λiλj]
/angbracketleftμiλi/angbracketright/angbracketleftμjλj/angbracketright(11)
/epsilon1−
i./epsilon1−
j=/angbracketleftλiλj/angbracketright[μiμj]
[λiμi][λjμj](12)
/epsilon1−
i./epsilon1+
j=/angbracketleftλiμj/angbracketright[μiλj]
[λiμi]/angbracketleftμjλj/angbracketright(13)
We also list for convenience
/epsilon1+
i.pj=/angbracketleftμiλj/angbracketright[λiλj]
/angbracketleftμiλi/angbracketright(14)
and
/epsilon1−
i.pj=/angbracketleftλiλj/angbracketright[μiλj]
[λiμi](15)
Evidently, in this formalism, flipping helicity corresponds to interchanging the brackets
/angbracketleft.../angbracketrightand [ ...].
We need one more important observation. Obviously, in a tree-level diagram for n-gluon
scattering, you cannot have as many 3-gluon vertices as you like. Draw the tree diagramsforn=4 for example (see figure N.2.3). The number of 3-gluon vertices could be either
0 or 2. In general, the number of 3-gluon vertices can be at most n−2. You are asked to
verify this in exercise N.2.2. As remarked earlier, while the 4-gluon vertex does not involvemomentum, the 3-gluon vertex is linear in momentum. Thus, in the numerator of theFeynman amplitude, we have npolarization vectors /epsilon1
ibut at most n−2 momenta. We are
to form a scalar out of these Lorentz vectors by taking dot products. Clearly, there are atleast two polarization vectors who have to dance with each other. Therefore we concludethat the tree amplitude must contain at least one power of /epsilon1
i./epsilon1j. (In the n=5 case that
power was actually 2, as we saw.)
Now we are ready to rock. For the amplitude A(++ ...+)(suppressing the momentum
labels), we simply choose the spinors μirepresenting the gauge degrees of freedom to all
be equal. Then all dot products /epsilon1+
i./epsilon1+
jbetween polarization vectors vanish according to
(11). But we just argued that the tree amplitude must contain at least one power of /epsilon1i./epsilon1j.
Remarkably, we have shown that the maximal helicity-violating amplitude vanishes for any
n! Our intuition suggested that these amplitudes are suppressed, but in fact they vanish.
What about the next-to-maximal helicity-violating amplitudes with one negative helicity,
namely A(−++ ...++)? Label the gluon with negative helicity as 1. Once again, for
i=2,...n, choose μiall equal to λ1. Then /epsilon1+
i./epsilon1+
j∝/angbracketleftμiμj/angbracketright=0, for i,j/negationslash=1. Furthermore,
/epsilon1−
1./epsilon1+
i∝/angbracketleftλ1μi/angbracketright=/angbracketleftλ1λ1/angbracketright=0 for i/negationslash=1. The amplitude A(−++ ...++)also vanishes!
Clearly, this “cheap” trick of exploiting gauge freedom no longer works for the next
amplitude with two negative helicities. T o see why the trick does not work any more, lookatA(−−+ ...++)for instance. Once again we could, for i=3,...n, choose μ
iall equal
N.2. Gluon Scattering | 491
(a)3 2
4 13 2
4 13 2
4 1
(b) (c)
Figure N.2.3
so that /epsilon1+
i./epsilon1+
j=0 fori,j≥3, but then we don’t have enough freedom to make all the other
polarization dot products vanish. In fact, at some point, we better have some nonvanishingamplitudes. In the literature, these amplitudes with two negative helicities are calledmaximal helicity-violating amplitudes. Upon crossing two of the gluons, they describetwo gluons producing n−2 gluons, with helicities + +→++ ...+,− +→−+ ...+,
and− −→−−+ ...+.
Explicit calculation of A(1,2,3,4)
Then=4 case is the simplest. T ake a deep breath and try to calculate A(1−,2−,3+,4+)
andA(1−,2+,3−,4+). For 4-gluon scattering these two are the only nonvanishing tree
amplitudes, since by parity the amplitudes with three minuses are related to the amplitudeswith three pluses (which we know vanish), and so on.
The bad news is that the calculation is fairly involved. The good news is that we can still
exploit gauge freedom mercilessly and that the final answer is surprisingly simple.
T ackle A(1
−,2−,3+,4+)first. The relevant diagrams are shown in figure N.2.3. Let us
simplify the notation as much as possible: write /angbracketleft12/angbracketright=/angbracketleftλ1λ2/angbracketright, [12] =[λ1λ2], and so forth.
Now we need the colored Feynman rules in the form given in appendix 1. In line with
the preceding discussion let us choose ˜μ1=˜μ2=˜λ3andμ3=μ4=λ2. Then all but one
of the polarization dot products vanish. For instance, /epsilon1−
2./epsilon1+
3∝/angbracketleftλ2μ3/angbracketright[μ2λ3]∝/angbracketleftλ2λ2/angbracketright=0.
The only nonzero product is /epsilon1−
1./epsilon1+
4=/angbracketleftλ1μ4/angbracketright[μ1λ4]/([λ1μ1]/angbracketleftμ4λ4/angbracketright)=/angbracketleft12/angbracketright[34]/([13]/angbracketleft24/angbracketright),
where the second equality follows from our gauge choice. This implies that the quarticdiagram N.2.3a vanishes, since it involves the product of two polarization dot products.
We notice that there are only two more diagrams (fig. N.2.3b,c) rather than three. With
the traditional Feynman rules there is a diagram with 1 and 3 on the same cubic vertex.Here we see another advantage of color stripping. We are looking at the coefficient oftr(T
a1Ta2Ta3Ta4). The diagram we just described has Ta1next to Ta3and so does not
contribute to this particular color ordering.
Next, the diagram in figure N.2.3b vanishes. Look at the cubic vertex involving 2, 3, and
v(for the virtual gluon): (/epsilon12./epsilon13/epsilon1v.p2+/epsilon13./epsilon1v/epsilon12.p3+/epsilon1v./epsilon12/epsilon13.pv), with /epsilon1vunderstood
as a “placeholder” to be contracted with the /epsilon1∗
vfrom the other cubic vertex. The first term
492 | Part N
vanishes because /epsilon12./epsilon13=0, the second term because /epsilon12.p3∝[μ2λ3]=[λ3λ3]=0, and the
third term because /epsilon13.pv=−/epsilon13.(p2+p3)=−/epsilon13.p2∝/angbracketleftμ3λ2/angbracketright=/angbracketleftλ2λ2/angbracketright=0. Our gauge
choice was wise indeed!
Only one diagram (fig. N.2.3c) left to calculate. The cubic vertex (/epsilon11./epsilon12/epsilon1v.p1+/epsilon12./epsilon1v/epsilon11.
p2+/epsilon1v./epsilon11/epsilon12.pv)is to be contracted with the other cubic vertex (/epsilon13./epsilon14/epsilon1∗
v.p3+/epsilon14./epsilon1∗
v/epsilon13.
p4+/epsilon1∗
v./epsilon13/epsilon14.(−pv)). In each of these vertices, the first term vanishes, since the only
nonzero polarization product is /epsilon11./epsilon14. T o obtain the amplitude we replace the polarization
product /epsilon1ρ
v/epsilon1ω∗
vfor the placeholder by the propagator −igρω/(p1+p2)2=−igρω/(2p1.p2).
Again, since all but one of the polarization dot products vanish, only one term survives
the contraction with gρω. We obtain A(1−,2−,3+,4+)=/epsilon11./epsilon14/epsilon12.p1/epsilon13.p4/p1.p2.
Since we are after conceptual understanding more than anything else, now and hence-
forth, in this and the next two chapters, we will suppress overall factors to keep variousexpressions as uncluttered as possible.
We have already calculated /epsilon1
1./epsilon14, so it remains to evaluate /epsilon12.p1=/angbracketleft21/angbracketright[31]/[23], /epsilon13.
p4=/angbracketleft24/angbracketright[34]//angbracketleft23/angbracketright, and p1.p2=/angbracketleft12/angbracketright[12]. Thus A=/angbracketleft12/angbracketright[34]2/([12][23] /angbracketleft23/angbracketright).
We can now use various identities to write this in a more symmetric form. First, momen-
tum conservation gives/summationtext
ip(i)
α˙α=/summationtext
iλ(i)
α˜λ(i)
˙α=0. Multiplying this by εβαε˙α˙γλ(j)
β˜λ(k)
˙γwe ob-
tain/summationtext
i/angbracketleftji/angbracketright[ik]=0 for any jandk. Second, we have /angbracketleft34/angbracketright[34]=p3.p4=p1.p2=/angbracketleft12/angbracketright[12].
Finally, the spinors can be regarded as 2-dimensional vectors and so any two spinors μand
νspan the space. Thus a third spinor λcan always be expanded as a linear combination of
the other two, viz, λ=(/angbracketleftλν/angbracketrightμ−/angbracketleftλμ/angbracketrightν)/ /angbracketleftμν/angbracketright, with the coefficients determined easily by
contracting with μandν. Contracting with a fourth spinor ηthen yields
/angbracketleftλη/angbracketright/angbracketleftμν/angbracketright=/angbracketleftλν/angbracketright/angbracketleftμη/angbracketright−/angbracketleftλμ/angbracketright/angbracketleftνη /angbracketright (16)
known as the Schouten identity.
Using these identities we now massage Ainto shape. Multiply the numerator and
denominator of Aby/angbracketleft34/angbracketrightto obtain /angbracketleft12/angbracketright2[34]/(/angbracketleft23/angbracketright/angbracketleft34/angbracketright[23]). Next, multiply the numerator
and denominator by /angbracketleft12/angbracketright2. In the denominator write /angbracketleft12/angbracketright[23]=− /angbracketleft 14/angbracketright[43]. Finally, we
obtain (suppressing overall phase factors, as promised)
A(1−,2−,3+,4+)=/angbracketleft12/angbracketright4
/angbracketleft12/angbracketright/angbracketleft23/angbracketright/angbracketleft34/angbracketright/angbracketleft41/angbracketright=p1.p2
p2.p3(17)
Compare this with figure N.2.2. You should be impressed, even though here we are doing
then=4 rather than the n=5 case.
Recall that we have another amplitude A(1−,2+,3−,4+)yet to calculate, in which the
two negative-helicity gluons are not adjacent in color. You should work this out as anexercise, but it turns out that we can use a trick. Write the analog of (2) for 4-gluon scattering
M=i/summationdisplay
permutationstr(Ta1Ta2Ta3Ta4)A(1−,2+,3−,4+) (18)
We have already remarked that if we extend the gauge group from SU(N) toU(N) , the
extra gluon (known in the literature, perhaps confusingly, as the “photon”) does not coupleto the other gluons (because the couplings in Y ang-Mills theory all involve commutators;see appendix 1.) Thus if we replace, say T
a2, by the identity matrix, the entire sum should
vanish. The six terms in the sum then break up into two groups, multipled by either
N.2. Gluon Scattering | 493
tr(Ta1Ta3Ta4)or tr(Ta1Ta4Ta3). Since the two traces are independent, the two groups
vanish separately. The traces in the three terms tr (Ta1Ta2Ta3Ta4)A(1−,2+,3−,4+)+
tr(Ta1Ta3Ta2Ta4)A(1−,3−,2+,4+)+tr(Ta1Ta3Ta4Ta2)A(1−,3−,4+,2+)all become
tr(Ta1Ta3Ta4). We thus obtain the so-called photon decoupling identity A(1−,2+,3−,4+)
+A(1−,3−,2+,4+)+A(1−,3−,4+,2+)=0, relating the desired amplitude to two ampli-
tudes already known from (17). Thus
A(1−,2+,3−,4+)=−(A(1−,3−,2+,4+)+A(1−,3−,4+,2+))
=− /angbracketleft 13/angbracketright4/parenleftbigg1
/angbracketleft13/angbracketright/angbracketleft32/angbracketright/angbracketleft24/angbracketright/angbracketleft41/angbracketright+1
/angbracketleft13/angbracketright/angbracketleft34/angbracketright/angbracketleft42/angbracketright/angbracketleft21/angbracketright/parenrightbigg
=/angbracketleft13/angbracketright4
/angbracketleft12/angbracketright/angbracketleft23/angbracketright/angbracketleft34/angbracketright/angbracketleft41/angbracketright(19)
where we used the Schouten identity.
Remarkably, the two amplitudes A(1−,2−,3+,4+)andA(1−,2+,3−,4+)have the same
form. It is tempting to conjecture that for n-gluon scattering, the maximal helicity-violating
amplitudes in which two of the gluons carry negative helicity and the rest positive helicityis given by the elegant expression (for n≥4)
A(1+,2+,...j−,... ,k−...n+)=/angbracketleftjk/angbracketright4
/angbracketleft12/angbracketright/angbracketleft23/angbracketright/angbracketleft34/angbracketright.../angbracketleft(n−1)n/angbracketright/angbracketleftn1 /angbracketright(20)
This conjecture was first put forward by Parke and T aylor and proved by Berends and Giele
(using an off-shell recursion method and a precursor to the on-shell recursion method tobe explained in the next chapter.) We will prove it in the next chapter.
Meanwhile, we note that one way of arguing for the conjecture’s validity is to verify
that the proposed amplitude satisfies all the symmetry requirements. Besides Lorentzinvariance (obviously satisfied), amplitudes at tree level in a massless theory like pureY ang-Mills should also satisfy scale and conformal invariance.
One interesting check is to count, for each i, the powers of λ
iminus the powers of ˜λi.
Call this quantity /Lambda1i. Then since momentum has the form ∼λ˜λ, it contributes 0 to /Lambda1i.I n
contrast, for negative helicity /epsilon1−
α˙α=λα˜μ˙α/[λ,μ]∼λ/˜λ. For positive helicity we have the
opposite: /epsilon1+
α˙α=μα˜λ˙α//angbracketleftμ ,λ/angbracketright∼˜λ/λ. Thus we have /Lambda1i=− 2hi.
We checked that indeed, in (20), we have /Lambda1i=2 fori=j,kand/Lambda1i=− 2 fori/negationslash=j,k.
Keeping track of /Lambda1iduring the calculation also provides us with a useful check.
Note that the n=5 scattering amplitude, which we started this chapter with, is
completely determined, since there are only two independent nonzero amplitudes:A(1
−,2−,3+,4+,5+)andA(1−,2+,3−,4+,5+).
Further developments
The astonishing simplicity of (20) has sparked a surge of interest and further develop-
ments. Here I will be content to mention some of them.
Once the tree amplitudes are done, one can calculate loop amplitudes by using a
more sophisticated version of the unitarity methods and of the Cutkosky cutting rulesof chapter II.8. Proceeding in this way, various authors have bootstrapped their way up to
494 | Part N
multiloop amplitudes. While the actual computational labor can quickly get out of hand,
it is still enormously less than the labor needed with traditional Feynman methods.
What about the basic cubic vertex of the theory? We will work it out in appendix 2 and
show that it fits nicely into the form in (20) with one important caveat.
Surely you, the astute reader, feel that there must be some deep reason for the aston-
ishing simplification from the mess in figure (1) to the elegant expression in (20). Indeed,tree amplitudes in gauge theories (and in gravity) turn out to be even simpler when writ-ten in terms of the twistors studied by Penrose decades ago. As this exciting development
2
occurred while this book was going to press, I have to balance my desire to make the bookas up-to-date as possible against pagination constraints. Thus I can provide here only anultra-concise (and hence perhaps somewhat cryptic) key to the literature, giving you nomore than a flavor of what is involved.
Include the momentum conservation delta function with the amplitudes of the type
studied here and define M(... ,λ
i,˜λi,...)≡A(λ ,˜λ)δ(4)(/summationtextn
j=1λj˜λj). Due to space con-
straints, I will suppress the kinematic dependence of Mon all but the particle iand write
simply M(λi,˜λi). Let us Fourier transform Min two possible ways (and overuse the letter
Msomewhat):
M(W i)=/integraldisplay
d2λiexp(i˜μα
iλiα)M(λ i,˜λi). (21a)
and
M(Z i)=/integraldisplay
d2˜λiexp(iμ˙α
i˜λi˙α)M(λ i,˜λi). (21b)
where W≡(˜μ,˜λ)andZ≡(λ,μ)denote two 4 −component objects which may be re-
garded for the time being as column “vectors.” The intent here is to transform Mse-
quentially for i=1, 2, ...nusing either (21a) or (21b). Consider SO( 2, 2)here instead of
SO( 3, 1), so that the spinors λand˜λare real, and hence we can take μand˜μto be real
as well. Thus, these integral transforms are no more and no less than the Fourier trans-forms you have long been familiar with, and the variable μis conjugate to the variable ˜λ
in the same sense that pis conjugate to qin quantum mechanics. The objects WandZ,
known as a twistor and a dual twistor and conjugate to each other, each consisting of 4real components, naturally transform under the group SL( 4,R)(namely the set of all 4 by
4 matrices with real entries and unit determinant), with the invariant W.Z=˜μλ+˜λμ.
Given more than one W’s and Z’s we also have the Lorentz invariants Z
1IZ2≡<λ1,λ2>
andW1IW2≡[λ1,λ2]. (Here I, in a slightly abused notation used in the literature, evidently
denotes the 4 by 4 matrix containing the 2 by 2 identity matrix either in its upper left corneror in its lower right corner depending on whether it acts on WorZ, with all other entries
equal to zero.)
We have (displaying the helicity hof particle iwhile suppressing the index i)M(tW ,h)=/integraltext
d
2λexp(it˜μλ)M(λ ,t˜λ,h)=t−2/integraltext
d2λ/primeexp(i˜μλ/prime)M(t−1λ/prime,t˜λ,h)=t2(h−1)M(W ,h)
where we used the observation earlier that /Lambda1=− 2h, namely that M(t−1λ,t˜λ)=
t2hM(λ ,˜λ). Similarly, M(tZ ,h)=t−2(h+1)M(Z ,h). This scaling result, which you realize
comes from the little group (see p. 186), indicates that we should favor a mixed or am-
2The literature on twistors could be traced starting with the paper mentioned on p. 477.
N.2. Gluon Scattering | 495
bitwistor representation for the scattering amplitude, using Wwhen the particle carries
+helicity and Zwhen the particle carries −helicity.
For example, for the basic Y ang-Mills cubic vertex with helicities (++− )(see appendix
2) we write M(W+
1,W+
2,Z−
3). The scaling relation just derived imposes powerful con-
straints on this amplitude, namely M(W1,W2,Z3)=M(tW 1,W2,Z3)=M(W 1,tW 2,Z3)=
M(W 1,W2,tZ3), which implies that in the ambitwistor representation the defining ver-
tex for Y ang-Mills theory is apparently, up to an irrelevant overall constant, just 1! Moreprecisely, M(W
1,W2,Z3)depends on the three possible invariants W1.Z3,W2.Z3, and
W1IW 2. The scaling relations (note that tcould be either positive or negative) then force
Mto have the amazingly simple form
M(W+
1,W+
2,Z−
3)=sign(W1.Z3)sign(W2.Z3)sign(W1IW2)
In different kinematic regions, the basic Y ang-Mills vertex is numerically equal to ±1.
T ree amplitudes live naturally in ambitwistor space. As another example, the 4-gluon
scattering amplitude (19) we worked hard to get becomes simply
M(W+
1,Z−
2,W+
3,Z−
4)=sign(W 1.Z2)sign(Z2.W3)sign(W 3.Z4)sign(Z4.W1)
Let’s anticipate a bit and write the basic cubic vertex for gravity to be given in (N.3.20)
in this ambitwistor representation. Indeed, the scaling relations derived above couldbe immediately applied to the graviton, for which h=± 2. We obtain M(tW ,++)=
t
2M(W ,++) andM(tZ ,−−)=t2M(Z ,−−), thus immediately fixing the cubic vertex
for gravity to be
M(W++
1,W++
2,Z−−
3)=|(W1.Z3)(W2.Z3)(W1IW2)|
Going from Y ang-Mills to Einstein-Hilbert, we merely have to replace the sign function by
the absolute value!
Clearly, the take-home message is that quantum field theory possesses hidden structures
that the traditional Feynman diagram approach would likely have no hope of uncovering.
Appendix 1: Colored Feynman rules for Yang-Mills theory
Using the double line formalism of chapter IV .5, we can draw the cubic and quartic vertices in Y ang-Mills theory
as in figure IV .5.2. Our conventions for the generators of SU(N) are [Ta,Tb]=ifabcTcand tr (TaTb)=1
2δab.
Thusfabc=− 2itr([Ta,Tb]Tc). Start with the Feynman rule for the quartic vertex given in chapter VII.1 and
appendix C. First, fabefcde=− 4tr([Ta,Tb][Tc,Td]). Next we multiply by polarization vectors and obtain the
colored rule for the quartic vertex:
4ig2tr(TaTbTcTd)(/epsilon1 1./epsilon12/epsilon13./epsilon14−/epsilon14./epsilon11/epsilon12./epsilon13) (22)
The two other terms are obtained by permutation. Similarly, the cubic vertex in (C.18) becomes (with a trivial
change k→p)
−4igtr(TaTbTc)(/epsilon1 1./epsilon12/epsilon13.p1+/epsilon12./epsilon13/epsilon11.p2+/epsilon13./epsilon11/epsilon12.p3) (23)
As described in the text, we can now strip off the color factors tr (TaTbTc)and tr (TaTbTcTd).
Color stripped amplitudes satisfy a number of useful identities. For example, the color stripped amplitude for
n-gluon scattering satisfies the reflection identity A(1, 2, ... ,n)=(−1)nA(n ,... ,2 ,1). T o show this, note that
the stripped quartic vertex (/epsilon11./epsilon12/epsilon13./epsilon14−/epsilon14./epsilon11/epsilon12./epsilon13)does not change sign under the reflection 1234 →4321,
while the stripped cubic vertex changes sign under 123 →321. From exercise N.2.2, V3+2V4=n−2, and thus
V3is odd or even according to whether nis odd or even.
496 | Part N
Appendix 2: The cubic vertex in the spinor helicity formalism
A rather natural question to ask is what the cubic vertex (23) looks like in the spinor helicity formalism.
The first observation is that if we put all momenta on shell, p2
1=p2
2=p2
3=0, then the cubic vertex actually
vanishes. By momentum conservation, we have p2
1=(p2+p3)2=p2.p3=0. The conditions pi.pj=0 then
imply all three lightlike momenta point in the same direction, so that pi=Ei(1, 0, 0, 1 ),i=1, 2, 3. But this
means that, for example, /epsilon13.p1∝/epsilon13.p3=0, and thus the cubic vertex (23) vanishes.
Now you see the motivation for allowing the momenta to be complex. Then the conditions pi.pj=0n o
longer force all three lightlike momenta to point in the same direction, and we can have a nonvanishing cubic
vertex on shell. As explained in the text, to complexify momentum, we simply remove the constraint ˜λ=λ∗.B y
the way, by complexifying the momenta here, we are anticipating the discussion in the next chapter a bit.
As always, we are free to choose the μspinors to our advantage. A good choice here is μ1=μ2andμ3=
λ1. Referring to (12) and (13), we then have /epsilon1−
1./epsilon1−
2∝[μ1μ2]=0 and /epsilon1−
1./epsilon1+
3∝/angbracketleftλ1μ3/angbracketright=0. The cubic vertex
collapses to
A(1−,2−,3+)=/epsilon1−
2./epsilon1+
3/epsilon1−
1.p2=/parenleftbigg/angbracketleftλ2μ3/angbracketright[μ2λ3]
[λ2μ2]/angbracketleftμ3λ3/angbracketright/parenrightbigg/parenleftbigg/angbracketleftλ1λ2/angbracketright[μ1λ2]
[λ1μ1]/parenrightbigg
=/angbracketleft12/angbracketright2
/angbracketleft13/angbracketright[μ1λ3]
[μ1λ1](24)
As in the text, we are ignoring all overall factors.
T o get rid of the unphysical μ1, we need a variant of the momentum conservation identity given in the text.
Multiplying/summationtext
ip(i)
α˙α=/summationtext
iλ(i)
α˜λ(i)
˙α=0b yεβαε˙α˙γλ(j)
β˜μ˙γ, we obtain/summationtext
i[μλi]/angbracketleftλiλj/angbracketright=0 for any j, which for j=2
implies [ μ1λ3]/angbracketleftλ3λ2/angbracketright=− [μ1λ1]/angbracketleftλ1λ2/angbracketright.
Multiplying (24) by /angbracketleftλ3λ2/angbracketright//angbracketleftλ3λ2/angbracketrightand applying the identity just derived we finally obtain the “mostly minus”
cubic vertex
A(1−,2−,3+)=/angbracketleft12/angbracketright4
/angbracketleft12/angbracketright/angbracketleft23/angbracketright/angbracketleft31/angbracketright(25)
Satisfyingly, we have obtained an expression consistent with (20) (which we have not yet proven) but keep in
mind that (25) holds only for complex momenta. I leave it to you to obtain the “mostly plus” cubic vertex
A(1+,2+,3−)=[12]4
[12][23][31](26)
which also follows from the rule about flipping helicities stated in the text.
What about the “all plus” and “all minus” vertices? By now you should be able to determine them as a simple
exercise.
Exercises
N.2.1 Work out the two polarization vectors for general μand˜μfor a gluon moving along the third direction.
N.2.2 Show that the number of cubic vertices in tree-level n-gluon scattering can be at most n−2.
N.2.3 Show that the result in (17) satisfies the reflection identity A(1−,2−,3+,4+)=A(4+,3+,2−,1−).
N.2.4 Show that the “all plus” and “all minus” cubic Y ang-Mills vertices (see appendix 2) vanish. [Hint: Choose
theμspinors wisely.]
N.2.5 Why doesn’t the argument in the text that A(−++ ...++)vanish apply to A(−++ )?
N.2.6 Insert the expression for the cubic vertex into (21) and derive M(W+
1,W+
2,Z−
3).
N.2.7 Show that M(W+
1,Z−
2,W+
3,Z−
4)reproduces (19).
N.2.8 Show that SL( 4,R)is locally isomorphic to the conformal group. [Hint: Identify the 15 =42−1 genera-
tors of the conformal group (3 rotations Ji, 3 boosts Ki, 1 dilation D, 4 translations Pμ, and 4 conformal
transformations Kμ) with the 15 traceless real 4 by 4 matrices.]
N.3 Subterranean Connections in Gauge Theories
Excess baggage
This text, like all texts on field theory, sings the praise of gauge theories—hey, Nature loves
them regardless of what physicists like—but, unlike many texts, emphasizes repeatedlythat gauge symmetry is strictly speaking not a symmetry, but a redundancy in description.Extra degrees of freedom are introduced only to be gauge fixed away. In the first edition ofthis book, I expressed in the Closing Words the hope that in the future physics will finda more elegant way of formulating this peculiar concept of local invariance. Perhaps thathope is being realized sooner rather than later!
In our current formulation of gauge theories, for a process involving nmassless gauge
bosons (photons or gluons) we are instructed to laboriously calculate an off-shell amplitudeM
μ1μ2...μn.
But experimentalists don’t know about amplitudes carrying Lorentz indices! SE from
chapter III.1 speaks up again. “My gauge bosons are specified by their helicities hi,i=
1,...n, not a Lorentz index.”
Come to think of it, we theorists do go through a strange two-step procedure involving
a lot of excess baggage. After toiling to obtain Mμ1μ2...μnwith external momenta off
shell, we then set external momenta on shell and contract with polarization vectors to
determine the scattering amplitude for gluons in specified polarization states Mλ1λ2...λn≡
/epsilon1λ1μ1/epsilon1λ2μ2.../epsilon1λnμnMμ1μ2...μn|onshell . In effect, in step 2 we wash away much of the unnecessary
information in Mμ1μ2...μnwe worked hard to get in step 1.
The cancellation in the 5-gluon scattering in the preceding chapter, with ∼10,000 terms
boiling down to a single term, should have convinced you that the traditional Feynmanway may not be so good. In your study of physics, you surely have had the pleasure ofwatching terms canceling against each other toward the end of a calculation, but 10,000
terms down to 1, that was the mother of all cancellations.
The key is that in gauge theories there is a kind of secret subterranean connection
between different Feynman diagrams, and cancellations are routine. Gauge invariance tells
498 | Part N
us that p1
μ1(/epsilon12
μ2.../epsilon1n
μnMμ1μ2...μn|onshell), for example, vanishes. Thus, the many diagrams
that go into Mμ1μ2...μnmust know about each other in some intricate way. (We saw a
glimpse of that way back in chapter II.7 when we proved gauge invariance.)
TheS-matrix reloaded
In spite of the tremendous difficulties lying ahead, I feel
thatS-matrix theory is far from dead and that . . . much
new interesting mathematics will be created by attemptingto formalize it.
—T . Regge
1
The garbage of the past often becomes the treasure of the
present (and vice versa).
—A. Polyakov2
As discussed in chapter III.8, back in the 1950s and 1960s, dispersion theorists3tried
to forge ahead by studying the analytic properties of various amplitudes as functionsof their external Lorentz invariants, namely the Mandelstam variables sandtfor 2-to-
2 scattering and q
2in our simple vacuum polarization example. But once one gets past
2-to-2 scattering, the analytic structure becomes unwieldy. The program failed and wasswept into the dustbin of physics history. (However, you might know that, through a ratherconvoluted process, this massive effort eventually gave birth to string theory.)
Remarkably, some features of this program are being revived. In particular, in this
chapter we will discuss the notion of complexifying physical variables. In an interestingtwist, it turns out to be better to complexify the external momenta (to be explained below)rather than invariants like sandt. A historical aside: Landau apparently suggested on one
occasion that it might be useful to consider complex momenta.
Consider the amplitude M(p
i,hi)for tree-level scattering of nmassless particles with
momentum and helicity (pi,hi),i=1 ,..., n, with p2
i=0. (For a gauge theory, we will
define the amplitude with the color factors already stripped away. Also, suppress the trivialmultiplicative coupling constant dependence and drop all such overall factors as we movealong.)
The novel idea is to pick two external momenta p
r,ps, complexifying them while
keeping them on shell and maintaining momentum conservation. We take all momentaas incoming. At this stage we can keep the discussion general and not even specify thetheory except to stipulate that it contains only massless particles. But to fix ideas, you canimagine a gauge theory. For some complex number z, replace p
randpsby
pr(z)=pr+zq and ps(z)=ps−zq (1)
1T . Regge, Publ. RIMS, Kyoto University, 12 suppl.: pp. 367–375, 1977.
2A. Polyakov, Gauge Fields and Strings, CRC, 1987, p. 1.
3See, for example, G. Barton, Dispersion T echniques in Field Theory, W . A. Benjamin, 1965.
N.3. Connections in Gauge Theories | 499
Rps(z)
Lpr(z)
PL(z)
Figure N.3.1
T o keep pr(z)2=0 and ps(z)2=0, we need q.pr,s=0 and q2=0, which is possible only
if we make qcomplex. T o be explicit, go to a frame in which pr+pshas only a time
component and use units so that the time component is equal to 2. Then
pr=(1, 0, 0, 1 ), ps=(1, 0, 0, −1), q=(0, 1,i,0) (2)
A technical aside. This is why I mentioned SO( 2, 2)in appendix B and in the preceding
chapter: with a (++− − )signature one could satisfy the on-shell constraint without having
a complex q. Here I will stick with the more physical SO( 3, 1)and consider complex
momenta as explained in the preceding chapter. Another side remark: As you will see,the discussion goes through for any spacetime dimension d≥4.
With this set up the scattering amplitude M(z)becomes an analytic function of z. Think
of the complex momentum zqflowing into the diagram with p
r(z), cruising through some
of the internal lines, and then flowing out with −ps(z). Let us turn on our pole detector.
At tree level, a pole can arise only from a propagator carrying momentum zq+.... Thus
the tree amplitude M(z)has only simple poles, coming from diagrams of the type shown
in figure N.3.1. Divide piinto two sets LandR, with those in Lflowing into a blob on
the left-hand side and those in Rflowing into a blob on the right-hand side. The two
blobs are connected by a single propagator carrying momentum PL(z), which by arbitrary
convention we choose to flow into blob L. The two blobs are themselves tree amplitudes in
the theory. Let nLandnRbe the number of external momenta in sets LandR, respectively
(withnL+nR=n, of course, and nL≥2,nR≥2). Then the left-hand blob represents tree
scattering of nL+1 particles, with nLparticles on shell and one particle with momentum
PL(z)off shell, with an amplitude ML(z). Similarly, the right-hand blob represents tree
scattering of nR+1 particles, with nRparticles on shell and one particle with momentum
−PL(z)off shell, with an amplitude MR(z).
500 | Part N
Clearly, the momentum PL(z) depends on zonly if pr(z) andps(z) do not appear in
the same set. With no loss of generality let pr(z)belong to the set Landps(z)to the set
R. Then PL(z)=−((/summationtext
i/epsilon1Lpi)+zq)=PL(0)−zqandPL(z)2=PL(0)2−2zq.PL(0)=
−2q.PL(0)(z−zL), where zL=PL(0)2/(2q.PL(0)). Thus Mhas a pole at z=zL, which,
sinceqis complex, is in general complex.
The amplitude M has poles all over the complex z-plane, at z=zL, one for each
valid partition of the external momenta into L+R. The residue has the factorized form
RL=ML(zL)MR(zL)/(2q.PL(0)), where ML(zL)andMR(zL)are now both on-shell
amplitudes, since the particle carrying momentum PL(zL)is now on shell. As always, we
suppress all inessential overall factors.
If, and that is a crucial if, M(z)→0a sz→∞ , then/contintegraltext
C(dz/z) M(z)=0, where the
contour Cis a circle of infinite radius running along infinity. We then shrink the contour,
picking up the pole at z=0, which contributes M(0)to the contour integral, and a bunch
of poles at z=zL, contributing a sum of terms consisting of the residue at each pole,
multiplied by 1 /zL. We thus determine the scattering amplitude to be
M(0)=−/summationdisplay
L,hRL
zL=−/summationdisplay
L,hML(zL)MR(zL)
PL(0)2(3)
Note that the sum also runs over the helicity hcarried by the intermediate particle PL.
The notation has been a bit compact, but suffices to get the essential point across
without cluttering the page with bloated expressions. But let us now make the notationa bit more precise. T o start with, z
Lof course depends on the specific partition Lthrough
the momentum PL. T o make sure you follow, let us describe ML(zL)more explicitly. It is
an on-shell amplitude with (nL+1)particles coming in, respectively carrying momentum
and helicity (pr(zL),hr),(pi,hi)fori/epsilon1L ,i/negationslash=r, and(PL(zL),h). Two of the momenta are
complex, namely pr(zL)andPL(zL). Let us emphasize that by construction PL(zL)2=0
and so all particles are on shell. Similarly, MR(zL)is an on-shell amplitude with (nR+1)
particles coming in, respectively carrying momentum and helicity (ps(zL),hs),(pi,hi)for
i/epsilon1R ,i/negationslash=s, and(−PL(zL),−h).
The crucial point is that, amazingly, as was discovered by Britto, Cachazo, Feng, and
Witten, we can determine the n-point tree amplitude M(z)in terms of lower point on-shell
tree amplitudes, specifically (3) as a sum over products of the (nL+1)-point amplitude
ML(zL)and(nR+1)-point amplitude MR(zL). Note that n−1≥nL+1≥3 (similarly
fornR+1), and thus by applying these so-called BCFW recursion relations repeatedly, we
can calculate any on-shell tree amplitude in Y ang-Mills theory and in gravity in terms of an
irreducible 3-point amplitude. Furthermore, in the primitive 3-point on-shell amplitude, allLorentz invariants constructed out of the momenta vanish, since p
i.pj=1
2(pi+pj)2=0.
T o determine the loop amplitudes, Bern, Dixon, and Kosower have generalized the
unitarity methods sketched earlier in this text and alluded to in the preceding chapter. With
these methods, one can calculate all amplitudes, trees and loops, and thus determine thetheory completely in terms of the helicity dependence of the 3-point on-shell amplitude.
N.3. Connections in Gauge Theories | 501
Amazingly, the old dream of the S-matrix school comes true! Everything within pertur-
bation theory is determined without our ever having to refer to a Lagrangian.
Note that to obtain physical amplitudes we need only M(z=0)but to recurse to higher
point amplitudes we need to know M(z/negationslash=0). As we will see, once we have M(z=0)we
can obtain M(z/negationslash=0)by analytic continuation.
As emphasized in the preceding chapter, the decomposition of a lightlike vector in terms
of spinors
pα˙α≡pμ(σμ)α˙α=λα˜λ˙α (4)
works equally well for complex lightlike vectors. In that case, as already explained in
chapter N.2, the two spinors λand˜λare independent of each other.
Another side remark: The deformation (1) considered here has a nice form in the
helicity spinor formalism of the preceding chapter. Let pr=λr˜λrandps=λs˜λs(with
spinor indices suppressed). Then the spinor deformation ˜λr→˜λr+z˜λsandλs→λs−zλr
(leaving λrand˜λsunchanged) gives the desired momentum deformation with q=λr˜λs,
which we see is not hermitean and hence corresponds to a complex momentum. This isconsistent with the discussion in the preceding chapter, since the deformation obviouslydoes not respect the equality between ˜λandλ
∗necessary for real momenta.
The naive person about to recurse
Imagine that you woke up one morning and had the wonderful idea of complexifying
momentum. Then suppose you had enough wits, after a bow to Cauchy, to discover thesemarvellous recursion relations. But after you calmed down, you wanted to try the recursionout on some theory. Naturally, you first chose a scalar field theory, say a ϕ
3or aϕ4theory.
Your enthusiasm dies immediately. In these theories, the basic vertex is just a number,
the coupling. For an n-point amplitude, there are always some Feynman diagrams in which
pr(z)andps(z)meet at one of the basic vertices, and the entire diagram does not even
depend on z. The crucial assumption that M(z)→0a sz→∞ is simply not true.
Most physicists might give up at this point, but suppose you were possessed of strength
of character and decided to take a look at Yang-Mills theory, thinking that, after all, it seemedmuch more fundamental than some dumb scalar field theory. But a quick look convincesyou that things are even worse. Consider the diagram in figure N.3.2a contributing to the n-
gluon amplitude. Put p
r(z)andps(z)as “far apart” as possible to maximize the number of
propagators between them. There are (n−3)propagators, contributing a factor of 1 /zn−3
to the amplitude as z→∞ . But alas, this is overwhelmed by (n−2)cubic vertices with
each vertex linear in momentum, thus contributing a factor of zn−2.
You have not yet included the polarization vectors, which for z=0 are given by /epsilon1−
r=
q,/epsilon1+
r=q∗and/epsilon1−
s=q∗,/epsilon1+
s=q(note that q↔q∗under r↔s, since the two momenta
502 | Part N
(a)
(b)
(c)ps(z) pr(z)
ps(z) pr(z)ps(z) pr(z)n/H110022
Figure N.3.2
N.3. Connections in Gauge Theories | 503
point in opposite directions). We also have to deform them to maintain their orthogonality
with the corresponding momentum vectors:
/epsilon1−
r(z)=q,/epsilon1+
r(z)=q∗+zps (5)
and
/epsilon1−
s(z)=q∗−zpr,/epsilon1+
s(z)=q (6)
You should check that all conditions are satisfied, for example, /epsilon1+
r(z).pr(z)=(q∗+
zps)(pr+zq)=z(q∗q+pspr)=0. [In the notation of (N.2.9, 10), the polarization vectors
here correspond to the choice μr(z)=μr(0)=λs,˜μr(z)=˜μr(0)=˜λs,μs(z)=μs(0)=
λr,˜μs(z)=˜μs(0)=˜λr. The first equal sign in each relation simply emphasizes that we
choose not to deform the μ’s and ˜μ’s.]
Note the peculiar asymmetry between randsafter deformation: in particular, two of the
polarization vectors, /epsilon1+
rand/epsilon1−
s, grow with zand thus worsen the large zbehavior. Putting
it all together and referring to (5) and (6), you would conclude [with the notation Mhrhs(z)]
that
M−+
naive(z)→zn−2
zn−3=z,M−−or++
naive(z)→z2,M+−
naive(z)→z3(7)
which most certainly do not →0.
A seemingly unimportant comment that will become important later: Of course, some
of the gluons other than randscould first interact among themselves as shown in fig-
ure N.3.2b. This merely reduces the effective nin the discussion above for those particular
diagrams, and we reach the same naive estimates.
Reality more benign than expectation
Reality turns out to be much more benign than our naive expectation! Actually, amplitudesin Yang-Mills theory behave better than amplitudes in scalar field theory, the opposite ofwhat we thought.
This fact is either astonishing or not so astonishing, depending on how jaded you are. I
have to admit that it sounds a bit less amazing after learning in the preceding chapter that∼10,000 terms can cancel down to a single term.
We can even concoct a heuristic physical argument. Go back to Yang-Mills theory and
call the particles gluons, as before. Apply crossing to gluon s, so that we have an incoming
gluon rwith a huge momentum p
r(z)∼zqin the large zlimit, emerging as a gluon with
the huge momentum −ps(z)∼zq. The other (n−2)gluons have fixed momenta and are
thus soft. We have a hard gluon blasting through a soft gluon background, something like
a high energy gamma ray blasting through a magnetic field, and thus we do not expectmuch scattering as z→∞ , and even less scattering that would flip the helicity of the hard
gluon. (The situation is conceptually similar to electron scattering in an external Coulombpotential, discussed in chapter II.6, except that here the field excitation being scattered isof the same type as the background field.)
504 | Part N
But not so fast! Even though you and I have studied physics for years, we haven’t built
up much intuition about complex momenta. At least I speak for myself. Alternatively, wecould go to SO( 2, 2)and deal with real momenta, but we haven’t much experience with
signature (++− − )spacetime, either.
Background field method
Nevertheless, the picture of a hard gluon blasting through a soft background turns out
to be helpful in guiding us toward an elegant formulation of the problem. We splitthe Yang-Mills gauge potential (which we write as Aon this occasion) into two pieces,
A(x)=A(x)+a(x) , a background potential Awithout high momentum components in
its Fourier transform and a fluctuating potential awith high momentum components. (You
would do exactly the same split when studying a laser beam passing through a laboratorymagnetic field.) To develop this so-called background field method (which is useful forother problems besides this one), it pays to use the differential form notation used inchapter IV .5.
We split the transformation law A→UAU
†+UdU†intoA→UAU†+UdU†and
a→UaU†. In other words, the background Atransforms like a Yang-Mills potential, while
the fluctuating atransforms like a matter field in the adjoint representation. Plugging into
the field strength F=dA+A2=d(A+a)+(A+a)2, we find Fequal to the sum of the
background field strength F=dA+A2and the 2-form da+Aa+aA+a2=(∂μaν+
[Aμ,aν]+aμaν)dxμdxν≡1
2(Dμaν−Dνaμ+[aμ,aν])dxμdxν. Switching back from math
to physics notation and defining the shorthand notation D[μaν]≡Dμaν−Dνaμ, we have
Fμν=Fμν+D[μaν]−i[aμ,aν]. Here Dμaν=∂μaν−i[Aμ,aν] is the covariant derivative
(with respect to the background potential A) of the adjoint field a.
Since we have only two hard gluons interacting with the soft background, it suffices to
expand the Yang-Mills Lagrangian to quadratic order in a:
L=−1
2g2trFμνFμν
=−1
2g2tr/parenleftBig
FμνFμν+D[μaν]D[μaν]+2FμνD[μaν]−2iFμν[aμ,aν]/parenrightBig
+O(a3) (8)
Since in the action we integrate Lover spacetime, we are effectively allowed to inte-
grate by parts. Thus the third term in the parenthesis, tr (FμνDμaν)=tr(Fμν(∂μaν−
i[Aμ,aν]))“=”t r ((DμFμν)aν). Since the background field satisfies the field equation
DμFμν=0, this term vanishes. (You should not be surprised that the term linear in a
in the action is linear in the field equation.)
Thus, to study the propagation of athrough the background A, we can focus on the
Lagrangian quadratic in a:Lquad=−(1/g2)tr/parenleftbig
(Dμaν−Dνaμ)Dμaν−iFμν[aμ,aν]/parenrightbig
As
always, we need to fix the gauge. Upon integration by parts we have
trDνaμDμaν=tr(DμaμDνaν+iFμν[aμ,aν]) (9)
N.3. Connections in Gauge Theories | 505
Note that, unlike ordinary derivatives, when the gauge derivatives DνandDμpass each
other, they produce the field strength Fμ. [Verify this! You might recall the more mathe-
matical form (IV .5.13).] Thus a convenient way of fixing the gauge is to add tr (DμaμDνaν)
so that the gauge-fixed Lagrangian becomes
Lquad=−1
g2tr(DμaνDμaν−2iFμν[aμ,aν]) (10)
[Incidentally, this parallels precisely what was done in (III.4.8) to obtain the Feynman gauge
withξ=1.]
Return now to our problem of studying the large-z behavior of the scattering amplitude
Mλρ. Recall from (7) that we obtain Mλρ→zn−2/zn−3=z. (Also recall that polarization
vectors have not yet been included, and they could multiply this behavior by z0,z1,o rz2.)
The culprit is the derivative in the cubic vertex ∼Aa∂a sitting inside the first term in
(10). In contrast, the ∼AaAa piece in the first term and the second term tr Fμν[aμ,aν]i n
(10) insert quartic vertices that do not grow with z.
The situation confronting us is now best discussed by the clever trick of renaming
indices. First, understand that Lorentz invariance is broken by the presence of the back-ground field A
μ, to be regarded as given and fixed. (This is the same as in chapter VI.2: the
presence of a background magnetic field means that parity and time reversal are broken.)But now suppose we simply relabel indices and write
Lquad=−1
g2tr(ηabDμaaDμab−2iFab[aa,ab]) (11)
where ηabis nothing but the humble Minkowski metric.
The first term by itself enjoys a hidden “enhanced Lorentz” symmetry: an SO( 3, 1)
transformation on the indices a,bleaves the Lagrangian invariant. We now exploit this
hidden symmetry. Since the leading behavior of Mabfor large zcomes from repeated
insertion of the cubic vertex ηabaaAμ∂μabcontained in the first term in (11), we conclude
that the leading behavior must be proportional to ηab.
In contrast, with one insertion of the quartic vertex from the second term in (11), we
decrease the power of zby one, since it does not contain a derivative on the field a.B u t
we also break the hidden “enhanced Lorentz” symmetry, since Fabis fixed. On the other
hand, there is an extra bit of information: we know that it is antisymmetric in (ab). [Note
that an insertion of the quartic vertex ηabaaAμAμabcontained in the first term in (11) also
decreases the power of zby one, but its contribution is proportional to ηab.]
Thus the hidden “enhanced Lorentz” symmetry tells us that the amplitude expanded in
powers of zmust have the form
Mab=(cz+...)ηab+Aab+1
zBab+... (12)
withcsome unknown constant. The only thing we know about the matrix Aabis that it is
antisymmetric in (ab). (I am following the notation in the literature. If you are confused
between this matrix Aand the background gauge potential A(x) you need to go back to
square 1.)
506 | Part N
We still have gauge invariance in the form pra(z)Mab(z)εsb(z)=0 and
εra(z)Mab(z)psb(z)=0, giving us valuable information. For example, looking up the form
pr(z)=pr+zq, we obtain qaMab(z)εsb(z)=−(1/z)p raMab(z)εsb(z), but since from (5)
/epsilon1−
r(z)=q, this means that /epsilon1−
ra(z)Mab(z)εsb(z)=−(1/z)p raMab(z)εsb(z).
Let us now look at the specific helicity combinations for which we had naive expectations
in (7). Recall that we expected M−+(z)→z. In fact, since /epsilon1+
s(z)=qandpr.q=0, we have
M−+(z)=/epsilon1−
ra(z)Mab(z)ε+
sb(z)=−1
zpra{(cz+...)ηab+Aab+1
zBab+...}qb
=−1
zpraAabqb+O(1
z2)→1
z(13)
This amplitude behaves better than naive expectation by two powers of (1/z)!
Next, M−−(z)→z2naively, but in fact
M−−(z)=/epsilon1−
ra(z)Mab(z)ε−
sb(z)=−1
zpra{(cz+...)ηab+Aab+1
zBab+...}(q∗
b−zprb)
=−1
z(praAabq∗
b+praBabprb)+O(1
z2)→1
z, (14)
three powers better than naive expectation. Similarly, M++(z)→1/z. Note that these
conclusions hold for any n. If you have the strength, you might want to witness the
cancellations by explicitly calculating the various M’s for low values of n.
But not all helicity amplitudes behave better than naive expectation. We finally come
toM+−(z), which →z3naively. Looking at (5) and (6) we already see trouble, since both
/epsilon1+
ra(z)andε−
sb(z)grow like z. Now we have
M+−(z)=/epsilon1+
ra(z)Mab(z)ε−
sb(z)=(q∗
a+zpsa){(cz+...)ηab+Aab+1
zBab+...}(q∗
b−zprb)
=−cps.prz3+O(z2)→z3(15)
Incidentally, note that our intuition about complex momenta is a bit shaky. The helicity-
conserving amplitude (+→+ )[namely M+−(z)by crossing; recall that Mwas defined
with all momenta going in] behaves worse than the (+→− )amplitude M++(z), the
(−→+ )amplitude M−−(z), and the (−→− )amplitude M−+(z). The polarization
vectors are continued for complex momentum in a nonsymmetric fashion.
Confusio suddenly speaks up! “You haven’t yet exploited the gauge invariance of the
background field,” he says.
We forgot that he often appears in the company of SE. Indeed, he is right. Very good—
Confusio did not become an assistant professor for nothing.
Indeed, let us look at the cubic vertex in figure N.3.2a more carefully: we have a hard
gluon carrying momentum zq+...scattering off a background gluon carrying some
small momentum pinto a hard gluon with momentum zq+.... The coupling comes
from the term tr ∂μaν[Aμ,aν] in the Lagrangian, and thus to leading order in zthe vertex is
proportional to zqμ.Aμ(p). According to exercise VII.1.1, we can choose a gauge in which
qμ.Aμ(p)=A2+i3(p)=0, known as the Chalmers-Siegel space cone gauge.
We should check to see if this is possible, but to streamline the exposition let us
merely do the abelian case. With Aμ(x)→Aμ(x)−∂μ/Lambda1(x), the desired gauge choice
N.3. Connections in Gauge Theories | 507
requires q.A(p)=iq.p/Lambda1(p), and thus we can solve for /Lambda1(p) as long as q.p/negationslash=0. While
q.pr,s=0 by construction, generically there is no reason for q.pito vanish for i/negationslash=r,s.
So we conclude that indeed we can get rid of the offending cubic vertex.
But not so fast! What about figure N.3.2c, in which all the soft gluons interact with
each other to form one single soft gluon carrying momentum/summationtext
i/negationslash=r,spi=−(pr+ps)?
Since q.(pr+ps)=0, we cannot set q.A(pr+ps)=0 and the diagram in figure N.3.2c
remains. Thus, even though we managed to get rid of the cubic vertices in figure N.3.2a, b,our previous conclusion about the large-z behavior of Mstill stands.
“Wait! What about the color factor?” Confusio yells. Let us look at the color structure
we stripped off. From figure IV .5.2b we see that the cubic vertex in figure N.3.2c requiresthat the two hard gluons be adjacent in color. It is easiest to explain the terminology byan example: a red-green gluon and a blue-yellow gluon are not adjacent in color, but theyare both adjacent to a red-yellow gluon (and to a blue-green gluon). Note that the couplingtrF
μν[aμ,aν] also requires that the two hard gluons be color adjacent.
Thus, if the two hard gluons are not color adjacent, the large-z behavior of Mis
somewhat better, since now c=0 and Aab=O(1/z). Then M−+→1/z2instead of 1 /z,
M+−→z2instead of z3, while M−−andM++are not improved. Confusio deserves credit
for his partial triumph, and perhaps eventually should be given tenure.
The bottom line is that, contrary to naive expectation, amplitudes in gauge theory behave
well enough for the BCFW recursion program to work. We don’t even mind that M+−
behaves badly; it suffices for the program that M−,any helicityvanishes for large z.I n
particular, in appendix 1 we will show how to complete the calculation started in thepreceding chapter.
As indicated earlier, once we determine the tree amplitudes, we can in principle obtain
all loop amplitudes by using unitarity. In this modern revival of the S-matrix spirit, we deal
with only on-shell amplitudes. The message here is that traditional Feynman diagramscarry around an enormous amount of unnecessary off-shell baggage. A dramatic exampleis furnished by this innocuous looking Feynman integral
/integraldisplayd4l
(2π)4lμlνlρlλ
l2(l−k)2(l−p)2(l−q)2(16)
which you can evaluate most conveniently using dimensional regularization. Try it. The in-
tegral looks similar to the integrals we did back in chapters III.6,7, but looks are deceptive.The answer, if printed on a page, is a total black smudge (see http://online.kitp.ucsb.edu/online/colloq/bern2/oh/05.html). After all, this integral is just one piece of a physical am-plitude and by itself does not possess any nice qualities, such as gauge invariance.
All possible Lorentz invariant theories
Remarkably, not only does BCFW recursion allow us to determine all n-point on-shell
amplitudes in terms of a primitive 3-point on-shell amplitude, it also restricts all possibletheories for which the recursion works. Let us sketch how this is possible. We anticipate
508 | Part N
here, as we will explain in the next chapter, that the recursion program works for massless
spin 2 as well as for spin 1 particles. Consider a 4-point on-shell amplitude M. The point is
that we are free to deform different pairs (r,s)to determine M. Suppose we pick (r,s)=
(1, 4). Then Mis the sum of two pieces, one with a pole in s=(p1+p2)2=(p3+p4)2and
another with a pole in t=(p1+p3)2=(p2+p4)2. But we could have also picked (1, 2)for
example. That the physical 4-point on-shell amplitude M(z=0)constructed in different
ways must agree imposes powerful self-consistency conditions on the primitive 3-pointon-shell amplitude.
Perhaps not surprisingly, for spin 2 massless particles, Einstein gravity is the only
possible theory, while for spin 1 massless particles, Yang-Mills gauge theory. Indeed, thisresult was proven long ago by Weinberg using rather general arguments. But it is stillinstructive to see how the same result emerges from a strikingly different formalism.These self-consistency conditions also allow one to explore and search for other possibletheories.
It is crucial that the primitive 3-point on-shell amplitude M
3is evaluated for complex
momenta, which allow more freedom than garden-variety everyday real lightlike momenta.(As already noted in appendix 1 to chapter N.2, the Yang-Mills cubic vertex vanishesfor real lightlike momenta.) I remind you again that for complex momenta p
i=λi˜λi,
the two spinors λiand˜λiare independent of each other. Recall from chapter N.2 that
/angbracketleftij/angbracketright=/angbracketleftλiλj/angbracketright≡εαβλiαλjβand [ij]=[λiλj]=[˜λi˜λj]≡ε˙α˙β˜λi˙α˜λj˙β. Also, pi.pj=/angbracketleftij/angbracketright[ij].
The on-mass shell conditions pi.pj=0 then become /angbracketleft12/angbracketright[12]=0,/angbracketleft23/angbracketright[23]=0, and
/angbracketleft31/angbracketright[31]=0. Apparently there are several possible solutions. For example, we could have
all three square brackets vanish with all three angled brackets nonzero, or we could havetwo square brackets vanish, say [12] =[23]=0, with /angbracketleft31/angbracketright=0. But there are only two
independent 2-component spinors, so three spinors cannot be linearly independent [take
˜λ
1∝(0, 1)and˜λ2∝(1,w), then the third spinor ˜λ3is necessarily a linear combination
of the other two]. Thus, [12] =0 and [23] =0 mean that ˜λ1∝˜λ2and˜λ2∝˜λ3, respectively,
which implies that ˜λ3∝˜λ1and [31] =0. Of course, the discussion can be repeated with
square and angled brackets interchanged. Thus we conclude that
either /angbracketleft12/angbracketright=/angbracketleft 23/angbracketright=/angbracketleft 31/angbracketright=0 or [12] =[23]=[31]=0 (17)
(For example, if [12] =[23]=[31]=0, then ˜λ2=α2˜λ1and˜λ3=α3˜λ1, and momentum
conservation/summationtext
ipi=/summationtext
iλi˜λi=0 implies λ1+α2λ2+α3λ3=0. The information here
is in the coefficients, since three 2-component spinors are always linearly dependent.)
Thus, depending on the helicities, either M3=MH(/angbracketleft12/angbracketright,/angbracketleft23/angbracketright,/angbracketleft31/angbracketright)orM3=
MA([12], [23], [31] ).
Recall from the preceding chapter that /Lambda1i=− 2hi, where /Lambda1icounts the powers of λi
minus the powers of ˜λi. But acting on MH, this just counts the powers of λi. Write
MH=/angbracketleft12/angbracketrightd3/angbracketleft23/angbracketrightd1/angbracketleft31/angbracketrightd2and solve for the unknown d’s using /Lambda11=d2+d3=− 2hi, etc.
Thend1=h1−h2−h3,d2=h2−h3−h1, andd3=h3−h1−h2. For example, suppose
N.3. Connections in Gauge Theories | 509
the theory contains spin 1 massless particles, with different varieties labeled by an index
awhose range we need not specify. Then we have, for example,
M3(1−
a,2−
b,3+
c)=fabc/parenleftbigg/angbracketleft12/angbracketright3
/angbracketleft23/angbracketright/angbracketleft31/angbracketright/parenrightbigg
(18)
since the helicities h1=h2=− 1 and h3=+ 1 imply that d1=d2=− 1 and d3=3. At this
stagefabcis some unknown coefficient that depends on the particle variety. As required,
we have two positive powers of λ1andλ2and two negative powers of λ3. This confirms
what we obtained in appendix 1 to the preceding chapter.
Several remarks follow.
1. From pi=λi˜λithe spinors λand˜λhave mass dimension1
2. Thus M3has mass dimension 1,
as expected (recall that the cubic coupling in gauge theory has the form ∼/epsilon1./epsilon1/epsilon1.p).
2. We obtain the 3-point amplitude
M3(1+
a,2+
b,3−
c)=fabc/parenleftBigg
[12]3
[23][31]/parenrightBigg
(19)
by flipping helicities, which, as we have learned in the preceding chapter, amounts to
interchanging the roles played by λand˜λ, so that it given by square instead of angled
brackets.
3. Note the power of the spinor helicity formalism. We can immediately generalize to higher
integer spin sby scaling the /Lambda1i’s and hence the d’s up by a factor of s. Thus we simply raise
the round parenthesis in M3(1−
a,2−
b,3+
c)to power s. The cubic vertex for spin 2 is thus
given by M3(1−−
a,2−−
b,3++
c)=fabc(/angbracketleft12/angbracketright3/(/angbracketleft23/angbracketright/angbracketleft31/angbracketright))2.
4. Interchanging 1and 2, we see that fabc=−fbacforsodd. Thus for sodd (s =1, for example),
we cannot have a theory with only one variety of particles. We are compelled to introducethe index a(and call it color!).
5. For seven (s =2, for example) we can get away with only one variety. Call it the graviton. The
coefficient f
abccan be omitted and one of the two basic cubic vertices for Einstein gravity
is simply given by
M3(1−−,2−−,3++)=/parenleftbigg/angbracketleft12/angbracketright3
/angbracketleft23/angbracketright/angbracketleft31/angbracketright/parenrightbigg2
(20)
(The other vertex is of course obtained by replacing angled brackets by square brackets.)
More on the 3-graviton vertex in appendix 2.
6. Check out the power of the self-consistency argument sketched above. Consider the 4-point
amplitude M(1a,2b,3c,4d)in a theory with a variety of spin 1 massless particles. Apply
the recursion to construct Mas the sum of an amplitude with an schannel pole, evidently
proportional to fabefcdewith an implicit sum over the label eof the intermediate particle,
and an amplitude with a tchannel pole proportional to facefbde. Requiring Mconstructed
with different choices of (r,s)in the recursion to be the same then gives the constraint
fabefcde+facefbde+fadefbce=0 (21)
510 | Part N
K/H11002j/H11002
2/H110013/H11001n/H11002
4/H110011/H11001
Figure N.3.3
But we recognize this as just the defining relation (B.19) for the generators of a Lie algebra
[Ta,Tb]=ifabcTcwritten out in the adjoint representation! The coefficients fabcthat appear
in the primitive 3-point on-shell amplitude are the structure constants of the algebra. If thisis too abstract for you, verify it for SU( 2).
The recursion program produces Einstein gravity and Yang-Mills theory as the unique
low energy theory for massless spin 2 and spin 1 particles, respectively, with sufficientlygood large-z behavior for the recursion relations to be valid. Of course, we also know that
in the Lagrangian formalism, the powerful constraints of local coordinate invariance andlocal gauge invariance fix the actions for Einstein gravity and Yang-Mills completely.
Appendix 1
Here, as promised, we use the recursion approach to prove the result conjectured in the preceding chapter, that
forn-gluon scattering, the maximal helicity-violating amplitude is given by
A(1+,2+,...j−,...,n−)=/angbracketleftjn/angbracketright4
/angbracketleft12/angbracketright/angbracketleft23/angbracketright/angbracketleft34/angbracketright.../angbracketleft(n−1)n/angbracketright/angbracketleftn1 /angbracketright(22)
(Using the cyclicity of the amplitude, we have with no loss of generality let gluon ncarry negative helicity.)
We take r=nands=1, and deform ˜λn→˜λn+z˜λ1andλ1→λ1−zλn(leaving λnand˜λ1unchanged), in
other words, pn→pn+zqandp1→p1−zqwithq=λn˜λ1. Write (3) as (see fig. N.3.3)
A(1+,2+,...j−,...,n−)=A3(ˆ1+,2+,ˆK−)An−1(−ˆK+,3+,...,j−,...,ˆn−)/PL(0)2(23)
Here we define K(z)=PL(z)to simplify writing. We use a hat to indicate that the corresponding momentum
has been complexified. Thus ˆ1,ˆn,ˆKremind us that p1(z),pn(z), and K(z)=−(p1(z)+p2)(evaluated at z=zL)
are the three complex momenta in the problem.
In the spirit of recursion, we are supposing that A3andAn−1 are given by (22) (and the corresponding
expression with all helicities flipped and angled brackets replaced by square brackets). Note that (22) does not
refer to any possible relation between the untwiddled λand twiddled ˜λspinors and thus makes sense for both
complex and real momenta. Notice that the sum over partition Land over hin (3) collapses to one term in (23).
We have used the result A(++ ...+)=0 and A(−+ ...+)=0, and the second half of (17) to eliminate a
diagram similar to that in fig. N.3.3 but with particles (n−1)andnparticipating in the cubic vertex instead of 1
and 2.
N.3. Connections in Gauge Theories | 511
Recursing and using PL(0)2=2p1.p2=2/angbracketleft12/angbracketright[12], we obtain [see (19)]
A(1+,2+,...j−,...,n−)=[ˆ12]3
[2,ˆK][ˆK,ˆ1]/angbracketleftjˆn/angbracketright4
/angbracketleftˆK3/angbracketright/angbracketleft34/angbracketright.../angbracketleft(n−1)ˆn/angbracketright/angbracketleftˆnˆK/angbracketright1
/angbracketleft12/angbracketright[12](24)
As before, we suppress overall constants.
The trick consists of taking various hats off or leaving them on. Since ˜λ1is unchanged, we can remove the
hat on ˆ1 when it appears in a square bracket. Similarly, since λnis unchanged, we can remove the hat on ˆnwhen
it appears in an angled bracket. On the other hand, we should leave the hat on ˆK. Instead, we use momentum
conservation −λK˜λK=λ1˜λ1+λ2˜λ2so that /angbracketleftˆK,3/angbracketright[ˆK,1 ]=− /angbracketleft 3,ˆK/angbracketright[ˆK,1 ]=/angbracketleft32/angbracketright[21] since [11] =0. (You might
note that this is the same sort of manipulation used to derive the first identity we needed to massage A4into
shape in the preceding chapter.) Similarly, /angbracketleftnˆK/angbracketright[2,ˆK]=− /angbracketleftnˆK/angbracketright[ˆK,2 ]=/angbracketleftn1/angbracketright[12].
Doing all this to (24) we obtain
A(1+,2+,...j−,...,n−)=[12]3/angbracketleftjn/angbracketright4
/angbracketleft12/angbracketright[12]/angbracketleftn1/angbracketright[12]/angbracketleft32/angbracketright[21]/angbracketleft34/angbracketright.../angbracketleft(n−1)n/angbracketright
=/angbracketleftjn/angbracketright4
/angbracketleft12/angbracketright/angbracketleft23/angbracketright/angbracketleft34/angbracketright.../angbracketleft(n−1)n/angbracketright/angbracketleftn1 /angbracketright, (25)
precisely the conjectured result.
Note how much more powerful the recursion approach is compared to the explicit spinor helicity calculation
we did to obtain A(1, 2, 3, 4 )in the preceding chapter, which in turn is so much more powerful than the traditional
Feynman diagram calculation. Thus theoretical physics marches on.
You might be puzzled that the quartic vertex (figure IV .5.1c) of Yang-Mills theory is not needed in the recursion
program. Does this nonparticipation in the program mean that we can multiply the quartic term in the Lagrangianby an arbitrary coefficient (including 0)? The resolution of this apparent paradox can be traced to the fact thatthe (perturbative) physical states of the gluon are built into the recursion relations. The quartic term is neededto guarantee gauge invariance and hence the two helicity states of the gluon.
Appendix 2
By showing you the mess in figure N.2.2 I have already plenty impressed upon you that the traditional Feynmandiagram approach is almost hopeless when it comes to gluons. The situation with gravity is far worse. Considerthe 3-graviton vertex. Conceptually it is easy to understand: we write g
μν=ημν+hμνand expand the Einstein-
Hilbert action (VIII.1.1) to O(h3). There it is, with indices suppressed, the cubic term h∂h∂h in (VIII.1.5). Of
course, this actually represents many terms with the eight indices contracted every which way, but which youcan readily work out. Next, pick the harmonic gauge for example, and derive the Feynman rule for the 3-gravitonvertex G
μα,νβ,σγ(p1,p2,p3), namely the analog of the 3-gluon vertex in (C.18). Each of the three gravitons, say
the one carrying momentum p1, can be created by any one of the three h’s in h∂h∂h, and thus many terms are
generated simply by permuting. The two derivatives give two powers of momentum. Thus, a typical term hasthe form p
1βp2μηανησγ.
Keep working! In all, Gμα,νβ,σγ(p1,p2,p3)contains about 100 terms. Now imagine calculating the one-loop
contribution to graviton-graviton scattering. You get the point.
By now, you fully appreciate that the traditional Feynman approach carries an enormous amount of unneces-
sary off-shell information. Already, if we put p1,p2, andp3on shell and contract Gμα,νβ,σγ(p1,p2,p3)with the
polarization vectors /epsilon1μα
1,/epsilon1νβ
2, and/epsilon1σγ
3, the 3-graviton vertex simplifies enormously to
G(p 1,p2,p3)=/epsilon1μα
1/epsilon1νβ
2/epsilon1σγ
3(p1σημν+cyclic)(p 1γηαβ+cyclic) (26)
Quite naturally, we can write the polarization vector for a spin 2 massless particle in terms of the polarization
vector for a spin 1 massless particle: /epsilon1μα(p)=/epsilon1μ(p)/epsilon1α(p). This form satisfies all that is required of a polarization
vector for spin 2: /epsilon1μα(p)p μ=0,/epsilon1μα(p)=/epsilon1αμ(p), and ημα/epsilon1μα(p)=0. Thus, indeed, the 3-graviton vertex
G(p 1,p2,p3)=[/epsilon1μ
1/epsilon1ν
2/epsilon1σ
3(p1σημν+cyclic)]2is the square of the 3-gluon vertex (N.2.23), in confirmation of (20),
which of course is just the same statement couched in another notation.
512 | Part N
Exercises
N.3.1 Show that the structure of Lie algebra (21) emerges naturally.
N.3.2 In appendix 1 we recursed by complexifying the momenta of two external lines with helicity +and−.
In the derivation of the recursion relation (3) we could have picked any two external lines to complexify.Determine the amplitude calculated directly in chapter N.2, namely A(1
−,2−,3+,4+), by complexifying
lines 1 and 2. This is an example of the self-consistency argument sketched in the text.
Particle physics experimentalists are fond of saying that yesterday’s spectacular discovery is today’s
calibration and tomorrow’s annoying background. The canonical example is the Nobel-winning discoveryof the CP-violating decay of the K
Lmeson into two pions. In theoretical physics, yesterday’s discovery is
today’s homework exercise and tomorrow’s trivium.
N.3.3 Using the explicit forms given for A(1−,2−,3+,4+)andA(1−,2+,3−,4+)in the preceding chapter,
check the estimated large zbehavior in (13–15).
N.3.4 Worry about the sloppy handling of factors of 2 in appendix 1. [Hint: The final result is correct because
the polarization vectors in (5–6) are normalized to |ε|2=2 for convenience.]
N.4Is Einstein Gravity Secretly
the Square of Yang-Mills Theory?
Gravity and gauge theory
Quantum gravity has baffled generations of theoretical physicists, as you have no doubt
heard. One aspect of this puzzle is the relationship between gravity and gauge theory,which describes the other fundamental interactions. While gravity and gauge theory areboth born of local invariance, the Einstein-Hilbert action/integraltext
d
4x√−gR and the Yang-Mills
action/integraltext
d4xtr(FμνFμν)look completely different.
Perturbatively, gravity is afflicted with an infinite number of interaction terms, as was
explained in chapter VIII.1, and hence gravity is not renormalizable, in stark contrast togauge theory. On the other hand, the two field theories enjoy many conceptual similaritiesbetween them. Yang-Mills theory is the unique low energy effective theory of a spin 1massless field, just as Einstein gravity is the unique low energy effective theory of a spin 2massless field.
String theory unifies gravity and gauge theory. This remarkable fact alone points to a
deep connection between gravity and gauge theory, even though within field theory theconnection is totally obscure. One important clue is that the oscillator spectrum of the openstring contains only the gauge field but not the graviton, which appears in the spectrumof the closed string. However, the closed string spectrum could be described as two copiesof an open string spectrum, thus leading Kawai, Lewellen, and Tye to discover relationsbetween graviton scattering and gauge boson scattering.
1In the limit of the string energy
scale going to infinity, we know that string theory reduces to field theory and thus a shadow
of these KLT relations should survive in field theory. (As you might know, not all theoristsare convinced that string theory corresponds to reality. If string theory eventually fails,its ultimate value might well turn out to be the light it sheds on the hidden structure ofquantum field theory.)
1It is definitely beyond the scope of this book to explain these statements. See, for example, J. Polchinski,
String Theory , p. 27.
514 | Part N
In any case, the bottom line is that string theory strongly hints that graviton ampli-
tudes can be expressed as products of Yang-Mills amplitudes, schematically Mgravitons ∼
Mgauge×Mgauge . The first reaction of many theoretical physicists when first told this
is puzzled skepticism. How is this possible, they ask quite reasonably, since Yang-Millscontains an internal symmetry group while gravity doesn’t?
Now that we have learned to strip color, a connection between amplitudes no longer
strikes us as so implausible, particularly if we stick to on-shell scattering amplitudes forgluons in specified polarization states M
λ1λ2...λn, namely amplitudes that experimentalists
can measure, rather than amplitudes Mμ1μ2...μncarrying Lorentz indices that theorists
using traditional methods play with. As we saw in the preceding chapter, the color strippedtree-level on-shell helicity amplitude for gauge boson scattering boils down to the /angbracketleft.../angbracketright
and [ ...] products of two component spinors. We did not do the analogous calculation of
the tree-level on-shell helicity amplitude for graviton scattering, but we could anticipatethat the result would again be expressed in terms of the /angbracketleft.../angbracketrightand [ ...] products. The
spinor helicity formalism is intrinsic to the Lorentz group SO( 3, 1), not tied to a specific
theory. In particular, the interaction vertices in Einstein gravity are again given in termsof scalar products of momenta and polarization vectors. Quite suggestively, the gravitonpolarization vectors can be written, as mentioned in appendix 1 to the preceding chapter,as/epsilon1
μν=/epsilon1μ/epsilon1ν, a product of the gauge theory polarization vectors.
Indeed, I have already given part of the mystery away in the preceding chapter. We saw
that the basic cubic interaction vertex of three gravitons (with complex momenta) is givenby the square of the corresponding quantity for three gluons.
In summary, thanks to our string theory friends, we now know that there exists a
secret structural connection between gravity and gauge theory that is totally opaque atthe Lagrangian level.
Deformed graviton polarizations
In this closing chapter, I give a brief introduction to the exciting quest for this secretconnection. I will be content to look at one specific calculation.
Go back to the BCFW recursion (chapter N.3). It would work for gravity if the com-
plexified scattering amplitude M(z)vanishes as z→∞ . But naively, it would seem that
the situation for gravity is even worse than the situation for gauge theory, since the cubicgraviton vertex is quadratic in momentum and thus goes like z
2. (Recall the two powers
of derivative in the scalar curvature; see chapter VIII.1.) Repeat the calculation in the pre-
ceding chapter for n-graviton on shell scattering. Go back to figure N.3.2 and interpret the
lines as gravitons. The (n−2)cubic vertices give a factor of z2(n−2)for large z, easily over-
whelming the factor of 1 /zn−3from the (n−3)propagators. This nasty behavior occurs
even before we include the polarization of the two hard gravitons.
The graviton carries helicity ±2 (appendix 2 of chapter VIII.1) and hence a polarization
“vector” /epsilon1μν, given by a symmetric and traceless tensor. We can naturally construct /epsilon1μν=
N.4. Gravity and Yang-Mills Theory | 515
/epsilon1μ/epsilon1νout of the polarization vectors for a massless spin 1 particle (as already explained in
the preceding chapter). Thus, after deformation,
/epsilon1++μν
r(z)=/epsilon1+μ
r/epsilon1+ν
r=(q∗+zps)μ(q∗+zps)ν,/epsilon1−−μν
r(z)=/epsilon1−μ
r/epsilon1−ν
r=qμqν(1)
and
/epsilon1++μν
s(z)=/epsilon1+μ
s/epsilon1+ν
s=qμqν,/epsilon1−−μν
s(z)=/epsilon1−μ
s/epsilon1−ν
s=(q∗−zpr)μ(q∗−zpr)ν(2)
Note that the /epsilon1μν(z)’s are in fact traceless and could go as either z0orz2for large z.
Putting it together, we obtain the naive estimate
M−−,++
naive(z)→z2(n−2)
zn−3=zn−1,M−−,−− or++,++
naive(z)→zn+1,M++,−−
naive(z)→zn+3(3)
The escalating behavior as nincreases is the hallmark of a nonrenormalizable theory, as
explained in chapter III.2.
Hard graviton in a soft spacetime
Once again, we hope that real life is cushier than naive expectation. By the same reasoningused for gauge theory, we study a hard graviton blasting through a gravitational field, thatis, a background of soft gravitons. So write the metric of spacetime as G
μν=gμν+hμν.
Plug this into the Einstein-Hilbert action (VIII.1.1) and extract the terms quadratic in h.
While the calculation is straightforward, it does involve some heavy lifting. To avoid thelabor, we note that using the harmonic gauge, we did this calculation in (VIII.1.10) butonly for the special case g
μν=ημν(in other words, we expanded around flat Minkowski
spacetime rather than a general curved spacetime). We had
L=1
64πG(ημνηλρηστ∂μhλσ∂νhρτ−1
2ημν∂μh∂νh) (4)
with the trace degree of freedom h≡ημνhμν. Henceforth, we set 64 πG=1.
Armed with symmetry considerations and our knowledge of gravity (chapter VIII.1), we
can almost immediately guess that when we go from a flat ημνto a curved gμνbackground,
this quadratic Lagrangian generalizes to
L=√−g(gμνgλρgστDμhλσDνhρτ−1
2gμνDμhDνh−2Rλρσ τhλσhρτ) (5)
withhnow defined as h≡gμνhμν. Here Ddenotes the covariant derivative with respect
to the curved metric gμνintroduced in chapter VIII.1 and Rλρσ τthe Riemann curvature
tensor constructed out of gμν. I trust you not to confuse this Dassociated with the
curved background with the covariant derivative in Yang-Mills theory used in the preceding
chapter and mentioned below in passing.
Let us go through the various features of (5). The√−ggoes with the spacetime volume
and is common to any Lagrangian in curved spacetime, as we learned way back in (I.11.2).We also learned there to promote any Lagrangian from flat to curved spacetime by replacingη
μνwithgμνand the ordinary derivative by a covariant derivative (see chapter IV .5). As you
516 | Part N
can see, everything pretty much works out in parallel with how things work out for gauge
theory. The new feature is the term involving the Riemann curvature tensor Rλρσ τ, which
vanishes upon restriction to flat spacetime. But you are not surprised that such a termcould pop up, given tr F
μν[aμ,aν] in (N.3.8). Indeed, the only thing we can’t determine
without doing the actual calculation is the numerical coefficient (−2)of this term. That
particular number will play no role in the following discussion.
Also, in the preceding chapter we dropped a term linear in aμbecause of the equation
of motion DμFμν=0. Here, analogously, we dropped a term linear in hμνbecause of
Einstein’s equation of motion Rμν=0. This also explains why terms involving the Ricci
tensor Rμνand the scalar curvature Rdo not appear in (5).
We want to calculate the large- zbehavior of various scattering amplitudes M−−,++,
etc., and compare with the naive expectation (3). We hope that the same trick we usedfor the gauge theory case would also work for gravity. Now the string theory hint, thatM
gravitons ∼Mgauge×Mgauge , suggests a factorized structure in the graviton amplitude,
and so more or less naturally leads to the guess that the first index λand the second index
σofhλσare somehow associated respectively with the two copies of Mgauge .
The key to breaking the problem apart is the Bern transformation unlinking these two
indices. Examining (5), we see that the only term that links the first index with the secondindex of h
λσappears in the term gμνDμhDνh, since h=gμνhμνdoes precisely that. How
to get rid of this term? The trick, following Bern and Grant, is to introduce a scalar field φ
and add the term 2 gμν∂μφ∂νφ. We are allowed to do this, since φdoes not appear in the tree
level graviton scattering amplitude we are studying. (Of course, the theory is changed frompure gravity, and φdoes circulate in loop diagrams for graviton scattering. Some readers
may also know that in string theory the graviton appears with a scalar φ, the experimentally
unobserved dilaton.)
For pedagogical clarity in explaining what we are going to do next, it is best to retreat
to the case of the flat background. Focus on the parenthesis in (4), now modified to(∂
μhλσ∂μhλσ−1
2∂μh∂μh+2∂μφ∂μφ). (The normalization of φis, in this context, just
chosen for convenience.) Since we can always make a field redefinition (see the appendix tochapter VIII.3) without affecting on-shell scattering amplitudes, we let h
λσ→hλσ+ηλσφ
(and hence h→h+4φ) andφ→φ+1
2h. You can verify that our parenthesis changes to
(∂μhλσ∂μhλσ−2∂μφ∂μφ). Since in this manipulation the role of ηλσis merely to convert
hλσintoh, the same transformation works when ηλσis promoted to gλσ.
The upshot is that we can effectively rewrite (5) as
L=√−g(gμνgλρgστDμhλσDνhρτ−2Rλρσ τhλσhρτ) (6)
Now that φhas done his job, we have unceremoniously thrown him out since he doesn’t
contribute to the on-shell tree amplitudes we are interested in. We have thus dropped thetermg
μν∂μφ∂μφ.
There has been quite a bit of formal development and perhaps the reader has lost sight
of what we are trying to do. Recall that we want to study the amplitude of a hard gravitonblasting through spacetime. Although a multitude of indices have appeared, as is alwaysthe case with gravity, you should recognize that this Lagrangian is conceptually simple: it
N.4. Gravity and Yang-Mills Theory | 517
is quadratic in the quantum field hdescribing the hard graviton and contains some given
c-number tensors gλρ(x)andRλρσ τ(x)pertaining to the background.
Unlinked melody
The important point is that the two indices carried by hλσare now unlinked from each
other in the first term in (6). In chapter VIII.1, we learned to trade a “world index” likeλfor a locally flat Lorentz index aby using the vierbein e
a
λ(x). Here we are invited to
introduce two sets of vierbein, eand˜e, with their associated connections ωand˜ω, and
writehλσ≡ea
λ˜e˜a
σha˜a. In reality, of course e=˜eandω=˜ω, but this notation keeps track of
the fact that the two sets of indices carried by hλσare unlinked. Note that hλσis treated in
our quadratic Lagrangian as just some tensor field living in a curved spacetime specifiedbyg
λσ≡ea
λ˜e˜a
σηa˜a.
Also in chapter VIII.1 we emphasized that the covariant derivative acting on vectors
carrying a world index and on vectors carrying a locally flat Lorentz index assumes differentforms, D
μVν=∂μVν−/Gamma1λ
μνVλandDμVa=∂μVa−ωb
μaVb, respectively. For pedagogical
clarity, I will use two different symbols DandDto denote what is conceptually the same
operation. For our problem we have Dλhμν=ea
μ˜e˜a
νDλha˜a, with Dλha˜a=∂λha˜a−ωb
λahb˜a−
˜ω˜b
λ˜aha˜b.
With this notation, the relevant Lagrangian becomes
L=√−g(gμνηab˜η˜a˜bDμha˜aDνhb˜b−2Rab˜a˜bha˜ahb˜b) (7)
We are now ready to study the large zbehavior of the scattering amplitude of a hard graviton
carrying momentum zq+...blasting through a curved background spacetime gμν.
The analysis proceeds much as in the Yang-Mills case discussed in the preceding chap-
ter. Focus on the first term:√−ggμνηab˜η˜a˜b(∂μha˜a−ωc
μahc˜a−˜ω˜c
μ˜aha˜c)(∂νhb˜b−ωd
νbhd˜b−
˜ω˜d
ν˜bhb˜d). The leading O(z2)behavior comes from the piece containing two derivatives in
the first term, namely Llead≡√−ggμνηab˜η˜a˜b∂μha˜a∂νhb˜b, and thus contributes to the am-
plitude a term proportional to ηab˜η˜a˜b. In the Yang-Mills case, the Lagrangian contains a
hidden “enhanced Lorentz” symmetry. Here the situation is even better: we have not one,but two hidden “enhanced Lorentz” symmetries. The term L
leadis evidently left invariant
by two separate SO( 3, 1)Lorentz transformations, one operating on the a,bindices, the
other on the ˜a,˜bindices.
The subleading O(z) behavior comes from the pieces in the first term containing one
derivative and one factor of either ωor˜ω, for example√−ggμνηab˜η˜a˜b∂μha˜a(ωd
νbhd˜b+
˜ω˜d
ν˜bhb˜d). In this way we find that Mab,˜a˜b→cz2ηab˜η˜a˜b+z(ηab˜A˜a˜b+Aab˜η˜a˜b)+..., with
Aaband˜A˜a˜btwo matrices antisymmetric in their indices. To see this, consider for example
the piece involving ω(after some relabeling of indices):√−ggμνηac˜η˜a˜b(∂μha˜a)ωb
νchb˜b. This
gives rise to the term Aab˜η˜a˜binMab,˜a˜b. Note that since the matrix Aabdepends on the
spin connection ωab
νof the background, all we can say is that it is antisymmetric in its two
indices ab. Recall that this is quite analogous to what we did in the Yang-Mills case.
518 | Part N
Before reading further, you could now flex your mental muscle and push ahead to
obtain the sub-subleading O(z0)behavior. This comes from the pieces in the first term
containing two factors of ωand˜ω, for example (again, after some relabeling of indices)√−ggμνηcd˜η˜a˜bωa
μcha˜aωb
νdhb˜b. All we can now conclude is that this contributes to the scat-
tering amplitude a term of the form Bab˜η˜a˜b, with Baban arbitrary matrix. To this order,
the second term in (7) also contributes, breaking the “enhanced Lorentz” symmetries com-pletely. Nevertheless, we can still exploit the known symmetry properties of the Riemanncurvature tensor under interchange of its indices to say something about its contributionto the scattering amplitude. Putting it all together we conclude that
Mab,˜a˜b→cz2ηab˜η˜a˜b+z(ηab˜A˜a˜b+Aab˜η˜a˜b)+Aab˜a˜b+(ηab˜B˜a˜b+Bab˜η˜a˜b)+O/parenleftbigg1
z/parenrightbigg
(8)
Compare this with (N.3.12), which states that the amplitude for the scattering of a hard
gluon off a background of soft gluons goes like Mab=(cz+...)ηab+Aab+1
zBab+....
Amazingly, you can see that the large- zbehavior for the scattering of a hard graviton off a
background of soft gravitons can be obtained by “squaring” the large-z behavior for the
scattering of a hard gluon off a background of soft gluons! In other words, Mab,˜a˜b∼
MabM˜a˜b, as far as the large-z behavior is concerned.
Just as in the Yang-Mills case, by exploiting gauge identities like pra(z)Ma˜a,b˜b/epsilon1sb˜b(z)=
0, we can determine the large- zbehavior of various helicity amplitudes. For your conve-
nience I remind you that (from the preceding chapter) pr(z)=pr+zqandps(z)=ps−
zq. Thus the gauge identity just displayed says that qaMa˜a,b˜b/epsilon1sb˜b(z)=
−(1/z)p raMa˜a,b˜b/epsilon1sb˜b(z). Recalling [see (1)] that /epsilon1−−μν
r(z)=qμqνwe see that in calcu-
lating the amplitude M−−,h(z)we can effectively replace /epsilon1−−μν
r(z)by(1/z2)pμ
rpν
r. Thus
we can immediately conclude, since /epsilon1++μν
s(z)∼z0for large z[see (2], that for example
M−−,++(z)→1
z2(9)
which is far better than the naive expectation M−−,++
naive(z)→zn−1. Indeed, the horrible
ever-escalating behavior with increasing nhas disappeared. Even more remarkably, the
large-z behavior of graviton scattering amplitudes is consistent with the string-inspired
notion that gravity is “the square of Yang-Mills.” Recall that in gauge theory M−+(z)→1/z.
Thus, for large z, indeed M−−,++(z)∼(M−+(z))2.
The bottom line here is that the large- zbehavior of gravity is surprisingly benign and
vanishes fast enough for the recursion program to work.
Gravity is a square?
So, is Einstein gravity secretly the square of Yang-Mills theory?
Already we have seen in the preceding chapter, anticipating that the recursion pro-
gram works, that the primitive 3-point amplitude for gravity (for one helicity configura-tion)(/angbracketleft12/angbracketright
3//angbracketleft23/angbracketright/angbracketleft31/angbracketright)2is the square of the primitive 3-point amplitude for gauge theory
N.4. Gravity and Yang-Mills Theory | 519
(/angbracketleft12/angbracketright3//angbracketleft23/angbracketright/angbracketleft31/angbracketright), something that you could have never suspected by staring at√−gR and
trFμνFμνuntill you are blue in the face.
The calculations in the previous section show that the large-z behavior for the scattering
of a hard graviton off a background of soft gravitons could be obtained by “squaring” thelargezbehavior for the scattering of a hard gluon off a background of soft gluons, certainly
something that nobody could have anticipated by looking at Lagrangians.
Further evidence that the answer to the title of this chapter is “yes” comes from a recent
calculation by Bern, Carrasco, and Johansson. Interestingly, they do not strip the colorfrom a Yang-Mills theory, but instead show that they can write the “color-dressed” treeamplitudes in the form
Atree(1, 2, ...,n)=/summationdisplay
anaca
(/Pi1jp2
j)a(10)
It is beyond the scope of this book to explain in detail how this expression is obtained. I
merely state that the index alabels an individual diagram. For each diagram, the amplitude
may be written as the product of a kinematic function naof the momenta and a color factor
ca, divided by the product of the momenta pjcarried by the internal lines. (I do not explain
here how naandcaare defined.) The tree amplitude is then given by a sum over all tree
diagrams.
Bern et al. then conjecture that the n-graviton scattering amplitude at the tree level is
given, amazingly, by
Mtree
gravity(1, 2, ...,n)=/summationdisplay
anana
(/Pi1jp2
j)a(11)
They have checked by explicit computation that their conjecture in fact holds up to n=8.
Furthermore, they have also verified that the conjecture, suitably generalized, also holdsfor the various supercousins of Einstein gravity and Yang-Mills theory.
Thus the evidence is extremely strong that, yes indeed, Einstein gravity is secretly the
square of Yang-Mills theory, at least at the level of tree amplitudes. However, as of thiswriting (February 2009), there is no definitive understanding within field theory. The finalword on the subject has yet to be said, and it is not even clear what the final path to the finalword might be. I would be foolish indeed to discuss this further in a textbook when theentire subject is being rapidly developed. By the time this book is published, the conjecturethat Einstein gravity is the square of Yang-Mills theory may well have been proved. If not,then nothing would please me more than if a reader of this textbook could go on and proveit, hopefully not just at the tree level, but to all orders.
What is the simplest field theory?
The uninitiated would likely answer ϕ4theory. Indeed, field theory texts almost all start with
some kind of scalar field theory. Even I am not able to do any better. But the sophisticated,namely you, now that you have reached the end of this text, realize that the more symmetrythe theory has the better. To theoretical physicists, simplicity actually secretly means
520 | Part N
symmetry. Incidentally, I have always hated scalar field theories, and have ventured to
say so publicly. It is hard to like the action L=1
2(∂ϕ)2−λϕ4, so barren of color and flavor.
Some of the major problems facing particle physics, such as the hierarchy problem, mayeventually turn out to stem from our not having mastered scalar field theory.
Of course, scalar field theory is the simplest in the superficial sense that you need to
know the least to approach it. As I said in chapter I.12, once one is familiar with scalarfield theory the rest consists of “merely” decorating the field with various indices describingspacetime or internal symmetries. But the symmetry and the resulting structure provideus with handles to grab on to. Both Yang-Mills theory and Einstein gravity have an internallogic sorely lacking in scalar field theory. As I mentioned in chapters VII.3 and VIII.4,the consensus view is that the first exactly soluble field theory would almost certainly beN=4 supersymmetric Yang-Mills theory, the supercousin of pure Yang-Mills theory. The
remarkable recent developments described in the last three chapters have only reinforcedthis view. Almost beyond belief, even gravity may be simpler than we had long thought.For large complexified momentum, graviton scattering for some helicity arrangementsactually behaves better than gluon scattering. The evidence is mounting that Einsteingravity may in some sense be the square of Yang-Mills theory. So now we are left with theamusing thought that the simplest field theory may well end up being gravity or N=8
supergravity with its maximal supersymmetry. (At this point, a friend of mine who workswithN=8 supergravity pipes up, “It sure doesn’t look simpler if you are the guy doing
the calculation!” It is clear from the simple dimensional argument of chapter III.2 thatas one goes to higher order, the numerator of the Feynman integrand quickly becomesextremely involved.)
Only time will tell who will win the simplest field theory contest, but we do have two
convincing candidates.
More Closing Words
In the closing words to the first edition of this book, I wrote that Yang-Mills theory
was almost begging for a better notation that would lay bare the deeper structure of thetheory. Oy, the excess baggage we have to carry! Ten thousand terms instead of one. Insome respects, the spinor helicity formalism and the recursion program explained inchapters N.2–N.4 provide a partial answer to that pious wish.
Imagine some theorist idly wondering, after 1865, if there were a better notation to
describe the six fields E
x,Ey,Ez,Bx,By,Bzfor which Maxwell had written 20 equations
(since he did not use vector notation). We can even fantasize that by fooling around withnumerology (“Look, 4 .3/2=6!”), this “crackpot” came up with an antisymmetric 4 by 4
matrix he called F. Shoehorning Maxwell’s equations in vacuum (some of them stating
that the time variation of EandBis related to the space variation of EandB) into this
strange notation, this guy could even stumble on a secret connection between space andtime.
The spinor helicity formalism and the recursion program, though elegant, are still
rooted in the perturbative expansion of the 1940s. Can they be pushed into the nonpertur-bative regime? There have been attempts in that direction.
In all previous revolutions in physics, a formerly cherished concept has to be jettisoned.
If we are poised before another conceptual shift, something else might have to go. Lorentzinvariance, perhaps? More likely, we may have to abandon strict locality. Again, in closingwords I mumble something (from steepest descent to integral to what?) about modifyingthe form of the path integral. The recursion program and the resuscitated S-matrix ap-
proach might be a step in this direction, formulating field theory while avoiding mention
of a local Lagrangian. But we need analyticity, and of course analyticity follows from local-ity and causality, as far as we understand. We know also that even local field theory couldspawn non-local constructs, most notoriously the horizon of a black hole. But there the
522 | Closing Words to Part N
dynamics bends the causal structure of spacetime out of whack. The lack of strict locality
is not built into the laws of physics.
Of course, we also know how to imbue physics with non-locality right from the start.
We have Wilson’s lattice formulation of gauge theory, and more recently, Wen’s intriguinglattice formulation of gravity.
When I showed the last three chapters to our friend SE, she mused, after some reflection,
“Now I see what theorists could always do when in doubt: enhance the symmetry and makeit local, complexify and bow to Cauchy, and take a square root when possible!”
I nodded, “These are the three ways of the warrior theorist: I call them the Einstein
way, the Heisenberg way, and the Dirac way. They were wildly successful in the past, andperhaps they will work in the future as well.”
With this edition of my textbook, I can no doubt count on a new group of readers to
come up with fresh insights into field theory. As these new chapters suggest, there maystill be plenty of secret structures to uncover. And thus field theory marches on.
Finally, I reveal the origin of the quote at the start of the preface to the second edition. As a
kid, Feynman came across a calculus book
1that proclaimed “What one fool can do, another
can.” He was thus inspired to master calculus. Now that you have mastered quantum field
theory, you can switch from the “understand” in the preface to the “do” in these closingwords.
1Silvanus P. Thompson (1851–1916), Calculus Made Easy , 1910, updated by Martin Gardner, St. Martin’s Press,
(1998). I am kind of trying to do for quantum field theory what Thompson did for calculus.
Appendix AGaussian Integration and the Central
Identity of Quantum Field Theory
The basic Gaussian:
/integraldisplay+∞
−∞dxe−1
2x2=√
2π (1)
The scaled Gaussian:
/integraldisplay+∞
−∞dxe−1
2ax2=/parenleftbigg2π
a/parenrightbigg1
2
(2)
Moments:
/integraldisplay+∞
−∞dxe−1
2ax2x2n=/parenleftbigg2π
a/parenrightbigg1
21
an(2n−1)(2n−3)...5.3.1,n≥1 (3)
Gaussian with source:
/integraldisplay+∞
−∞dxe−1
2ax2+Jx=/parenleftbigg2π
a/parenrightbigg1
2
eJ2/2a(4)
/integraldisplay+∞
−∞dxe−1
2ax2+iJx=/parenleftbigg2π
a/parenrightbigg1
2
e−J2/2a(5)
/integraldisplay+∞
−∞dxe1
2iax2+iJx=/parenleftbigg2πi
a/parenrightbigg1
2
e−iJ2/2a(6)
/integraldisplay+∞
−∞/integraldisplay+∞
−∞.../integraldisplay+∞
−∞dx1dx2...dxNei
2x.A.x+iJ.x=/parenleftbigg(2πi)N
det[A]/parenrightbigg1
2
e−(i/2)J.A−1.J(7)
/integraldisplay+∞
−∞/integraldisplay+∞
−∞.../integraldisplay+∞
−∞dx1dx2...dxNe−1
2x.A.x+J.x=/parenleftbigg(2π)N
det[A]/parenrightbigg1
2
e1
2J.A−1.J(8)
In what follows, we omit an overall factor.
Central identity of quantum field theory:
/integraldisplay
Dϕe−1
2ϕ.K.ϕ−V( ϕ) +J.ϕ=e−V( δ / δJ)e1
2J.K−1.J(9)
524 | Appendix A. Gaussian Integration
A trivial variation:
/integraldisplay
Dϕe−1
2ϕ.K.ϕ+J.ϕ=e1
2J.K−1.J(10)
Variations:
/integraldisplay
Dϕe(i/2)ϕ.K.ϕ+iJ.ϕ=e−(i/2)J.K−1.J(11)
/integraldisplay
Dϕei/integraltext
ddx[1
2ϕ(x)Kϕ(x)+J(x)ϕ(x) ]=ei/integraltext
ddx[−1
2J(x)K−1J(x) ](12)
/integraldisplay
Dϕe−/integraltext
ddx[1
2ϕ(x)Kϕ(x)+J(x)ϕ(x) ]=e/integraltext
ddx[1
2J(x)K−1J(x) ](13)
(where KorK−1or both may be nonlocal)
A specific example:
/integraldisplay
Dϕei/integraltext
ddx[(λ/2)ϕ2+ϕ¯ψψ]=ei/integraltext
ddx[−(1/2λ)(¯ψψ)2](14)
ForKhermitean with ϕcomplex:
/integraldisplay
Dϕ†Dϕe−ϕ†.K.ϕ+J†.ϕ+ϕ†.J=eJ†.K−1.J(15)
As noted earlier, various numerical factors have been swept under the integration measure. In applying these
formulas, be sure that these factors are not relevant for your purposes.
Appendix B A Brief Review of Group Theory
I give here a brief review of the group theory I will need in the text. I assume that you have been exposed to some
group theory, otherwise this instant review might not be intelligible. Most of the concepts are illustrated withexamples, and it goes without saying that you should work out all the examples and verify the assertions madewithout proof.
SO(N)
The special orthogonal group SO(N) consists of all NbyNreal matrices Othat are orthogonal
OTO=1 (1)
and have unit determinant
detO=1 (2)
We denote the element in the ith row and jth column by Oij. The group SO(N) consists of rotations in N-
dimensional Euclidean space and its defining or fundamental representation is given by the Ncomponent vector
/vectorv={vj,j=1 ,..., N}, which transforms under the action of the group element Oaccording to (as always, all
repeated indices are summed over)
vi→v/primei=Oijvj(3)
We define tensors as objects that transform as if they are equal to the product of vectors. For example, the tensor
Tijktransforms according to
Tijk→T/primeijk=OilOjmOknTlmn(4)
as if it is equal to the product vivjvk. The emphasis is on the phrase “as if”: Tijkis not to be thought of as being
equal to vivjvk.
It is important to develop some “feel” or intuition for groups and their representations. Some people find
it helpful to picture a certain number of objects being acted upon by the group and transformed into linearcombinations of each other. Thus, picture T
ijkasN3objects being scrambled together.
Tensors furnish representations of the group. In our particular example, each group element is represented
by an N3byN3matrix acting on the N3objects Tijk. The number of objects in a tensor is called the dimension
of the representation.
It may well be that any given object in a representation does not transform, under all the elements of the group,
into a linear combination of all the other objects, but only into a subset of them. Let me illustrate with an example.
526 | Appendix B. Brief Review of Group Theory
Consider Tij→T/primeij=OilOjmTlm. Form the symmetric Sij≡1
2(Tij+Tji)and antisymmetric combinations
Aij≡1
2(Tij−Tji). The symmetric combination Sijtransforms into OilOjmSlm, which is obviously symmetric.
Similarly, Aijtransforms into OilOjmAlm, which is obviously antisymmetric . In other words, the set of N2
objects contained in Tijsplit into two sets:1
2N(N+1)objects contained in Sijand1
2N(N−1)objects contained
inAij. TheSij’s transform among themselves and the Aij’s transform among themselves.
The representation furnished by Tijis said to be reducible: It breaks apart into two representations. Obviously,
representations that do not break apart are called irreducible.
We just exploited the obvious fact that the symmetry properties of a tensor under permutation of its indices is
not changed by the group transformation, namely that the indices on a tensor transform independently, as in (4).The various possible symmetry properties may be classified with Young tableaux, which is useful in a generaltreatment of group theory. Fortunately, in the field theory literature one rarely encounters a tensor with suchcomplex symmetry properties that one has to learn about Young tableaux.
Another way of saying this is that we can restrict our attention to tensors with definite symmetry properties
under permutation of their indices. In our specific example, we can always take T
ijto be either symmetric or
antisymmetric under the exchange of iandj.
We have yet to use the properties (1) and (2). Given a symmetric tensor Tijconsider the combination
T≡δijTij, known as the trace. Then T→δijT/primeij=δijOilOjmTlm=(OT)liδijOjmTlm=δlmTlm=T, where
we used (1). In other words, Ttransforms into itself. We can subtract the trace from Tijforming the traceless
tensor Qij≡Tij−(1/N)δijT. The1
2N(N+1)−1 objects contained in Qijtransform among themselves.
To summarize, given two vectors vandw, we can form a tensor, and decompose the tensor into a symmetric
traceless combination, a trace, and an antisymmetric tensor. This process is written as
N⊗N=[1
2N(N+1)−1]⊕1⊕1
2N(N−1) (5)
In particular, for SO( 3),3⊗3=5⊕1⊕3, a relation you should be familiar with from courses on mechanics
and electromagnetism.
There are two conventions for naming representations. We can simply give the dimension of the representa-
tion. (This can occasionally be ambiguous: Two distinct representations may happen to have the same dimension.)Alternatively, we can specify the symmetry properties of the tensor furnishing the representation. For instance,the representation furnished by a totally antisymmetric tensor of nindices is often denoted by [ n] and the rep-
resentation furnished by a totally symmetric traceless tensor of nindices by {n}. Obviously, [1] ={1}. In this
notation, the decomposition in (5) can be written as {1}⊗{ 1}={ 2}⊕{ 0}⊕[2]. For the group SO( 3), with its
long standing in physics, the confusion over names is almost worse than in reading Russian novels: For instance,{1}is also known as pand{2}asd.
We have yet to use (2). Using the antisymmetric symbol ε
123. . . N, we write (2) as
εi1i2...iNOi11Oi22...OiNN=1 (6)
or equivalently
εi1i2...iNOi1j1Oi2j2...OiNjN=εj1j2...jN (7)
By multiplying (7) by OTrepeatedly, we can obviously generate more identities. Instead of drowning in a sea of
indices, let me explain this point by specializing to say N=3. Thus, multiplying (7) by (OT)jNkN, we obtain
εi1i2i3Oi1j1Oi2j2=εj1j2j3(OT)j3i3
Speaking loosely, we can think of moving some of the O’s on the left hand side of (7) to the right hand side,
where they become OT’s.
Using these identities, you can easily show that [ n] is equivalent to [N −n], that is, these two representations
transform in the same way. For example, as is well known, in SO( 3)the antisymmetric 2-index tensor is equivalent
to the vector. (The cross product of two vectors is a vector.)
Any orthogonal matrix can be written as O=eA. The conditions (1) and (2) imply that Ais real and
antisymmetric, so that Amay be expressed as a linear combination of N(N−1)/2 antisymmetric matrices
denoted by iJij:O=eiθijJij(with repeated indices summed over). We have defined Jijas imaginary and
antisymmetric and hence hermitean. Since the commutator [ Jij,Jkl] is antihermitean, it can be written as a
linear combination of the iJ’s.
Ironically, some students are confused at this point because of their familiarity with SO( 3), which has special
properties that do not generalize to SO(N).
In speaking about rotations in 3-dimensional space we can specify a rotation as either around say the third
axis, with the corresponding generator J3, or as in the (1-2)-plane, with the corresponding generator J12=−J21.
Appendix B. Brief Review of Group Theory | 527
In higher dimensions, for example 10-dimensional space, we can speak of a rotation in the (6-7)-plane, with the
corresponding generator J67=−J76, but it is nonsense to speak of a rotation around the fifth axis. Thus, to
generalize to higher dimensions we should write the standard commutation relation [ J1,J2]=iJ3forSO( 3)as
[J23,J31]=iJ12, which can be generalized immediately to
[Jij,Jkl]=i(δikJjl−δjkJil+δjlJik−δilJjk) (8)
The right hand side reflects the antisymmetric character of Jij=−Jji. A potential confusion some students
may have about the notation: Jijdenotes a matrix generating rotation in the (i -j)-plane, a matrix with element
(Jij)klin the k-th row and l-th column. The indices i,j,k, andlall run from 1 to N, but in (Jij)klthe set {ij}
and the set {kl}should be distinguished conceptually: The former labels the generator and the latter are matricial
indices when the generator is regarded as a matrix. As an exercise, write down (Jij)klexplicitly and obtain (8) by
direct computation.
In studying group theory, as I have already remarked, one source of confusion comes from the fact that
some of the smaller groups, which we tend to encounter first in our studies, have special properties that do notgeneralize. The special property of SO( 3)we just noted is due to the fact that the antisymmetric symbol ε
ijkcarries
three indices and thus Jijmay be written as Jk≡1
2εijkJij.F o rSO( 4)the antisymmetric symbol εijklcarries four
indices and we can form the combinations1
2(Jij±1
2εijklJkl). Define J1
±≡1
2(J23±J14),J2
±≡1
2(J31±J24), and
J3
±≡1
2(J12±J34). By explicit computation, show that [ Ji
+,Jj
+]=iεijkJk
+,[Ji
−,Jj
−]=iεijkJk
−, and [ Ji
+,Jj
−]=0.
This proves the well-known theorem that SO( 4)is locally isomorphic to SO( 3)⊗SO( 3).
I assume that you know that SO( 3)is locally isomorphic to SU( 2). If you don’t, I give a brief review below.
With a few i’s included here and there, these two results prove the statement that the Lorentz group SO( 3, 1)
is locally isomorphic to SU( 2)⊗SU( 2), which we proved explicitly in chapter II.3. The Lorentz group can be
thought of as an “analytic continuation” of the rotation group SO( 4). See below for a more precise statement.
One highly non-obvious result of group theory is that SO(N) contains representations other than vector and
tensor. I develop the relevant group theory for the spinor representations in chapter VII.7.
SU(N)
We next turn to the special unitary group SU(N) consisting of all NbyNmatrices Uthat are unitary
U†U=1 (9)
and have unit determinant
detU=1 (10)
The story of SU(N) has more or less the same plot as the story of SO(N) with the crucial difference that the
tensors of the unitary groups can carry both upper and lower indices. We denote the element in the ith row and
jth column by Ui
j; the wisdom of this notation will soon become apparent.
The defining or fundamental representation of SU(N) consists of Nobjects ϕj,j=1 ,..., N, that transform
under the action of the group element Uaccording to
ϕi→ϕ/primei=Ui
jϕj(11)
Taking the complex conjugate of (11) we have
ϕ∗i→(Ui
j)∗ϕ∗j=(U†)j
iϕ∗j(12)
We invite ourselves to define an object we write as ϕithat transforms in the same way as ϕ∗i; thus
ϕi→ϕ/prime
i=(U†)j
iϕj (13)
Note that we did not say that ϕiis equal to ϕ∗i; we merely said that ϕiandϕ∗itransform in the same way.
As before, we can have tensors. The tensor ϕij
k, for example, transforms as if it is equal to the product ϕiϕjϕk:
ϕij
k→ϕ/primeij
k=Ui
lUj
m(U†)nkϕlm
n(14)
528 | Appendix B. Brief Review of Group Theory
Again, we emphasize that we did not say that ϕij
kis equal to ϕiϕjϕk. (In some books ϕiis called a covariant vector
andϕia contravariant vector. A tensor ϕ......... withmupper indices and nlower indices is defined to transform as
if it is equal to the product of mcovariant vectors and ncontravariant vectors.)
The possibility of complex conjugation in SU(N) leads naturally to having indices “upstairs” and “downstairs.”
Note that (9) can be written out explicitly as (U†)k
iUj
k=δj
iand thus the Kronecker delta in SU(N) carries one
upper and one lower index. It is important when taking traces that we set an upper index equal to a lower index
and sum over them: for example, we can consider δk
jϕij
k≡ϕij
j, which transforms as
ϕij
j→Ui
lUj
m(U†)njϕlm
n=Ui
lϕlm
m(15)
where we have used (9). In other words, ϕij
j, the trace of ϕij
k, denote Nobjects that transform into linear
combinations of each other in the same way as ϕi. Thus, given a tensor, we can always subtract out its trace.
As in the discussion for SO(N), tensors furnish representations of the group. The discussion proceeds as
before. The symmetry properties of a tensor under permutation of its indices are not changed by the grouptransformation.
Another way of saying this is that given a tensor we can always take it to have definite symmetry properties
under permutation of its upper indices and under permutation of its lower indices. In our specific example, we
can always take ϕij
kto be either symmetric or antisymmetric under the exchange of iandjand to be traceless.
Thus, the symmetric traceless tensor ϕij
kfurnishes a representation with dimension1
2N2(N+1)−Nand the
antisymmetric traceless tensor ϕij
ka representation with dimension1
2N2(N−1)−N.
Thus, in summary, the irreducible representations of SU(N) are realized by traceless tensors with definite
symmetry properties under permutation of indices. For example, in SU( 5), some commonly encountered
representations are ϕi,ϕij(antisymmetric), ϕij(symmetric), ϕi
j,ϕij
k(antisymmetric in the upper indices and
traceless) with dimensions 5, 10, 15, 24, and 45, respectively. Convince yourself that for SU(N) the dimensions
of the representations defined by these tensors are N,N(N−1)/2,N(N+1)/2,N2−1, and1
2N2(N−1)−N,
respectively.
The representation defined by the traceless tensor ϕi
jis known as the adjoint representation. By definition,
it transforms according to ϕi
j→ϕ/primei
j=Ui
l(U†)njϕl
n=Ui
lϕl
n(U†)n
j. We are thus invited to regard ϕi
jas a matrix
transforming according to
ϕ→ϕ/prime=UϕU†(16)
Note that if ϕis hermitean it stays hermitean, and thus we can take ϕto be a hermitean traceless matrix. (If ϕ
is antihermitean we can always multiply it by i.)Another way of saying this is that given a hermitean traceless
matrix X,UXU†is also hermitean and traceless if Uis an element of SU(N) .
As in the SO(N) story, representations of SU(N) have many names. For example, we can refer to the
representation furnished by a tensor with mupper and nlower indices as (m,n). Alternatively, we can refer
to them by their dimensions, with an asterisk to distinguish representations with mostly lower indices from therepresentations with mostly upper indices. For example, an alias for (1, 0)isNand for (0, 1)isN
∗. A square
bracket is used to indicate that the indices are antisymmetric and a curly bracket indicate that the indices aresymmetric. Thus, the 10 of SU( 5)is also known as [2, 0] =[2], where as indicated the 0 (no lower index) is
suppressed. Similarly, 10
∗is also known as [0, 2] =[2]∗.
The condition (10) can be written as either
εi1i2...iNUi1
1Ui2
2...UiN
N=1 (17)
or
εi1i2...iNU1
i1U2
i2...UN
iN=1 (18)
Thus, we have two antisymmetric symbols εi1i2...iNandεi1i2...iNthat we can use to raise and lower indices. Again,
we can immediately generalize (17) to
εi1i2...iNUi1
j1Ui2
j2...UiN
jN=εj1j2...jN
and multiplying this identity by (U†)jNpNand summing over jNwe obtain
εi1i2...pNUi1
j1Ui2
j2...UiN−1
jN−1=εj1j2...jN(U†)jNpN
Appendix B. Brief Review of Group Theory | 529
Clearly, by repeating this process, we can peel off the U’s on the left hand side and put them back as U†’s on the
right hand side. We can play a similar game with (18).
To avoid drowning in a sea of indices, let me show you how to raise and lower indices in a specific example
rather than in general. Consider the tensor ϕij
kinSU( 4). We expect that the tensor ϕkpq≡ϕij
kεijpq will transform
as a tensor with three lower indices. Indeed,
ϕkpq≡ϕij
kεijpq→εijpqUi
lUj
m(U†)nkϕlm
n=εlmst(U†)s
p(U†)tq(U†)nkϕlm
n=(U†)n
k(U†)sp(U†)tqϕnst
As in SO(N) we can look at the generators of SU(N) by noting that any unitary matrix can be written as
U=eiH, with Hhermitean and traceless as required by (9) and (10). There are (N2−1)linearly independent
NbyNhermitean traceless matrices Ta(a=1 ,2 ,..., N2−1). Any NbyNhermitean traceless matrix can be
written as a linear combination of the Ta’s and thus we can write U=eiθaTa, where θaare real numbers and the
index ais summed over.
Since the commutator [ Ta,Tb] is antihermitean and traceless, it can also be written as a linear combination
of the Ta’s:
[Ta,Tb]=ifabcTc(19)
(with the index csummed over.) The commutation relations (19) define the Lie algebra of SU(N), and fabcare
known as the structure constants. For SU( 2)the structure constants fabcare simply given by the antisymmetric
symbol εabc.
Sometimes students are confused by how the generators act. Consider an infinitesimal transformation
U/similarequal1+iθaTa. On the defining representation, ϕi→Ui
jϕj/similarequalϕi+iθa(Ta)i
jϕj. Thus, the ath generator acting
on the defining representation gives Taϕ. Now consider the adjoint representation (16)
ϕ→ϕ/prime/similarequal(1+iθaTa)ϕ(1+iθaTa)†/similarequalϕ+iθaTaϕ−ϕiθaTa=ϕ+iθa[Ta,ϕ] (20)
In other words, the ath generator acting on the adjoint representation gives [ Ta,ϕ]. Perhaps some students are
confused by the fact that ϕis used as a generic symbol to denote different objects.
Since the adjoint representation ϕis hermitean and traceless it can also be written as a linear combination of
the generators, thus ϕ=ϕbTb. Using (19) we can thus also write (20) as ϕc→ϕ/primec/similarequalϕc−fabcθaϕb. In particular
forSU( 2), the three objects ϕatransform as a 3-vector. (Note the notation: ϕais not to be confused with ϕi:i n
SU( 2)the index a=1, 2, 3 while i=1, 2.)
This last remark essentially amounts to a proof that SU( 2)is locally isomorphic to SO( 3). I will now give a
somewhat more formal proof. Any 2 by 2 hermitean traceless matrix Xcan be written as a linear combination of
the three Pauli matrices X=/vectorx./vectorσwith three real coefficients (x1,x2,x3), which we regard as the components of a
3-vector /vectorx. For any element UofSU( 2),X/prime≡U†XU is hermitean and traceless, so that we can write X/prime=/vectorx/prime./vectorσ.
Note that we have implicitly used the first defining property of an SU( 2)matrix (9). By explicit computation,
we find det X=−/vectorx2. Invoking the second defining property of an SU( 2)matrix (10), we obtain det X/prime=detX
and thus /vectorx/prime2=/vectorx2. The 3-vector /vectorxis rotated into the 3-vector /vectorx/prime. Thus we can associate a rotation with any given
U. Since Uand−U are associated with the same rotation, this gives a double covering of SO( 3)bySU( 2).A
physicist would just say that when a spin1
2particle is rotated through 2 π, its wave function changes sign. The
map clearly preserves group multiplication: if two elements U1andU2ofSU( 2)are mapped to the rotations R1
andR2respectively, then the element U1U2is mapped to the rotation R1R2. Alternatively, noting that tr X2=/vectorx2
and tr X/prime2=trX2, we obtain the same conclusion.
Once again, the two special unitary groups that most students learn first, namely SU( 2)andSU( 3), have
special properties that do not generalize to SU(N), just as SO( 3)has special properties that do not generalize to
SO(N) , possibly leading to confusion.
ForSU( 2), because the antisymmetric symbol εijandεijcarry two indices, it suffices to consider only tensors
with upper indices, all symmetrized: We can raise all lower indices of any tensor by contracting with εijrepeatedly.
After this is done, we can remove any pair of indices in which the tensor is antisymmetric by contracting withε
ij.
In particular, ϕi=εijϕj, which can be stated equivalently in terms of a special property of the Pauli matrices
σ2σ∗
aσ2=−σa (21)
530 | Appendix B. Brief Review of Group Theory
so that
σ2(ei/vectorθ/vectorσ)∗σ2=ei/vectorθ/vectorσ(22)
ForSU( 2)(11) becomes
ϕi→ϕ/primei=(ei/vectorθ/vectorσ)i
jϕj
Complex conjugating, we obtain
ϕ∗i→[(ei/vectorθ/vectorσ)i
j]∗ϕ∗j=[(−iσ 2)ei/vectorθ/vectorσ(iσ 2)]i
jϕ∗j
and so
iσ2ϕ∗→ei/vectorθ/vectorσ(iσ2ϕ∗)
We learn that iσ2ϕ∗transform in the same way as ϕ. Recall that we define ϕito transform in the same way as ϕ∗i.
Thus, εijϕjtransforms in the same way as ϕi. In the jargon, SU( 2)is said to have only real and pseudoreal
representations, but not complex representations. A pseudoreal representation is equivalent to its complexconjugate upon a similarity transformation. Recall that (21) figures into our discussion of charge conjugation inchapter II.1 and of the Higgs doublet in chapter VII.2.
ForSU( 3)it suffices to consider only tensors with all their upper indices symmetrized and all their lower
indices symmetrized. Thus, the representations of SU( 3)are uniquely labeled by two integers (m,n), where m
andndenote the number of upper and lower indices. The reason is that the antisymmetric symbols ε
ijkand
εijkcarry three indices. We can always trade a pair of lower indices in which the tensor is antisymmetric for one
upper index, and similarly for upper indices.
You can see easily that these special properties do not generalize beyond SU( 2)andSU( 3).
Multiplying representations together
In a course on quantum mechanics you learn how to combine angular momentum. We have already encountered
this concept in (5), which when specialized to SO( 3), tells us that 3 ⊗3=5⊕1⊕3, as we noted. This is
sometimes described by saying that when we combine two angular momentum L=1 states we obtain L=0, 1, 2.
Students are justifiably confused when this procedure is also known as addition of angular momentum.
Given two tensors ϕandηofSU(N) , with mupper and nlower indices and with m/primeupper and n/primelower indices,
respectively, we can consider a tensor Twith(m+m/prime)upper and (n+n/prime)lower indices that transforms in the
same way as the product ϕη. We can then reduce Tby the various operations described above. This operation of
multiplying two representations together is of course of fundamental importance in physics. In quantum fieldtheory, for example, we multiply fields together to construct the Lagrangian.
As an example, multiply 5
∗and 10 in SU( 5). To reduce Tij
k=ϕkηijwe separate out the trace ϕkηkj(which
transforms as a 5 )after which there is nothing more we can do. Thus,
5∗⊗10=5⊕45 (23)
As another example, consider 10 ⊗10:ϕijηkl. It is easiest to write ηklequivalently as a tensor with three lower
indices εmnhklηkl. The product 10 ⊗10 then carries two upper and three lower indices and we will write it as Tij
mnh.
Taking traces, we separate out Tij
mij, which we recognize as 5∗, and the traceless part of Tij
mnj, which we recognize
as 45∗(see above), thus obtaining:
10⊗10=5∗⊕45∗⊕50∗(24)
As exercises you can work out
5⊗5=10⊕15 (25)
and
5⊗5∗=1⊕24 (26)
You should recognize the 24 as the adjoint.
Appendix B. Brief Review of Group Theory | 531
In physics we are often called upon to multiply a tensor by itself. Statistics then plays a role. For instance,
SU( 5)grand unification contains a scalar field ϕitransforming as 5. Because of Bose statistics, the product ϕiϕj
contains only the 15.
Restriction to subgroup
To explain the next group theoretic concept, let me take a physical example. The SU(3) of Gell-Mann and Ne’eman
transforms the three quarks u,d, andsinto linear combinations of each other. It contains as a subgroup the
isospin SU( 2)of Heisenberg, which transforms uandd, but leaves salone. In other words, upon restriction to
the subgroup SU( 2)the irreducible representation 3 of SU( 3)decomposes as
3→2⊕1 (27)
Consider an irreducible representation with dimension dof some group G. When we restrict our attention
to a subgroup H, the set of dobjects will in general decompose into nsubsets, containing d1,d2,...,d nobjects,
such that the objects of each subset only transform among themselves under the action of H. This makes obvious
sense since there are fewer transformations in Hthan in G.
The decomposition of the fundamental or defining representation specifies how the subgroup His embedded
inG. Since all representations may be built up as products of the fundamental representation, once we know
how the fundamental representation decomposes, we know how all representations decompose. For example,inSU( 3)
3⊗3∗=8⊕1 (28)
while in SU( 2)
(2⊕1)⊗(2⊕1)=(3⊕1)⊕2⊕2⊕1 (29)
Comparing (28) and (29) we learn that
8→3⊕1⊕2⊕2. (30)
Alternatively, we can simply look at the tensors involved. Consider ϕiofSU( 3)where the index itakes on the
value 1, 2, 3. Let the index μtakes on the value 1, 2. Obviously, ϕi={ϕμ,ϕ3}corresponds to an explicit display
of (27). Then ϕi
j={¯ϕμ
ν,ϕμ
3,ϕ3
μ,ϕ3
3}, where the bar on ¯ϕμ
νis to remind us that it is traceless. This corresponds
precisely to (30).
Actually, SU( 3)also contains the larger subgroup SU( 2)⊗U(1), where the U(1)is generated by the traceless
hermitean matrix
⎛
⎜⎜⎝−100
0−10
00 2⎞
⎟⎟⎠
We can then write (27) as 3 →(2,−1)⊕(1, 2), where the notation is almost self-explanatory. Thus, (2,−1)
denotes a 2 under SU( 2)with “charge” −1 under U(1).
In the text, we will decompose various representations of SU( 5)andSO( 10). Everything we do there will
simply be somewhat more elaborate versions of what we did here.
More on SO(4),SO(3,1), and SO(2,2)
In chapter II.3 you learned that acting on the two objects ψαwithα=1, 2 in the spinor representation (1
2,0),
the generators of rotation and boost are represented by Ji=1
2σiandiKi=1
2σi, respectively. I remind you that
the equal sign means “represented by.” For most purposes (for example, classifying quantum fields) and at thelevel of rigor of this book, it suffices to think of the Lie algebra generated by commuting J
iandKi. Occasionally,
however, it is useful to contemplate the actual group with group elements ei/vectorθ/vectorJandei/vectorϕ/vectorK.
In the spinor representation (1
2,0)the group elements are represented by ei/vectorθ/vectorσ
2ande/vectorϕ/vectorσ
2. While ei/vectorθ/vectorσ
2is special
unitary, the 2 by 2 matrix e/vectorϕ/vectorσ
2, bereft of the i, is merely special but not unitary. (Incidentally, to verify these and
subsequent statements, since you understand rotation thoroughly, you could, without loss of generality, choose
532 | Appendix B. Brief Review of Group Theory
/vectorϕto point along the third axis, in which case e/vectorϕ/vectorσ
2is diagonal with elements eϕ
2ande−ϕ
2. Thus, while the matrix
is not unitary, its determinant is manifestly equal to 1.) This set of matrices defines the multiplicative groupSL( 2,C), consisting of all 2 by 2 complex-valued matrices with unit determinant.
Let us count the number of generators of this group. Two conditions on the determinant (real part =1,
imaginary part =0) cut the four complex entries containing eight real numbers down to six numbers, which
accounts for the six generators of the Lorentz group SO( 3, 1).
To exhibit the map explicitly, we extend the earlier discussion showing that SU( 2)covers SO( 3). Consider the
most general 2 by 2 hermitean matrix
XM=x0I−/vectorx./vectorσ=/parenleftBiggx0−x3x1−ix2
x1+ix2x0+x3/parenrightBigg
(31)
By explicit computation, det XM=(x0)2−/vectorx2. (To see this instantly, choose /vectorxto point along the third axis
and invoke rotational invariance.) Now consider X/prime
M=L†XML, with Lan element of SL( 2,C). Manifestly,
detX/prime
M=detXMand thus the transformation preserves (x0)2−/vectorx2and hence corresponds to Lorentz transfor-
mations. Since Land−L give the same transformation x→x/prime, we see that SL( 2,C)double covers SO( 3, 1).
Mathematicians say that SO( 3, 1)=SL( 2,C)/Z 2.I fL is also unitary, then x0/prime=x0and the transformation is
a rotation. The SU( 2)subgroup of SL( 2,C)double covers the rotation subgroup SO( 3)of the Lorentz group
SO( 3, 1), that is, SO( 3)=SU( 2)/Z 2.
Incidentally, if we introduce an iat a strategic location and define the 2 by 2 matrix XE=x4I+i/vectorx./vectorσ, regarding
(/vectorx,x4)as a 4-dimensional vector, we have det XE=(x4)2+/vectorx2, the Euclidean length squared of the 4-vector.
(Once again, choose /vectorxto point along the third axis so that XEis a diagonal matrix with elements x4±ix3.) Since
ei/vectorθ/vectorσ
2=cosθ
2+isinθ
2(ˆθ.σ)withˆθa unit vector in the θdirection (to see this, once again choose /vectorθto point along
the 3rdaxis), we see that XE/((x4)2+/vectorx2)1
2is an element of SU( 2). (We will come back to this observation in the
next section.) Thus, for any two elements UandVofSU( 2), the matrix X/prime
E=V†XEUcan also be decomposed
in the form X/prime
E=x/prime4I+i/vectorx/prime./vectorσ. Evidently, det X/prime
E=detXE. Thus the transformation preserves (x4)2+/vectorx2and
describes an element of SO( 4). This shows explicitly that SO( 4)is locally isomorphic to SU( 2)⊗SU( 2).I f
V=U, we have a rotation, and if V†=U, the Euclidean analog of a boost.
Note that while the rotation group SO( 3)is compact, the Lorentz group SO( 3, 1)is not, since the range of
the boost parameters /vectorϕis unbounded. In contrast, the group SO( 4)is compact and thus can be covered by a
compact group, namely, SU( 2)⊗SU( 2), but the noncompact group SO( 3, 1)cannot be.
At this point, having done SO( 4)andSO( 3, 1), I might as well (with a wink toward the nuts who complained
that this book is not encyclopedic enough) throw in the group SO( 2, 2)for use in part N. Let us strip the Pauli
matrix σ2(kind of a “troublemaker” or at least an odd man out) of his iand define (just for this paragraph)
σ2≡/parenleftBigg0−1
10/parenrightBigg
Any real 2 by 2 matrix XHcould be decomposed as XH=x4I+/vectorx./vectorσ. Now det XH=(x4)2+(x2)2−(x3)2−(x1)2,
the quadratic form of a spacetime with two time and two space coordinates. The set of all linear transformations(with unit determinant) on (x
1,x2,x3,x4)that preserve this quadratic form defines the group SO( 2, 2).
Introduce the multiplicative group SL( 2,R)consisting of all 2 by 2 real-valued matrices with unit determinant.
For any two elements LlandLrof this group, consider the transformation X/prime
H=LlXHLr. Evidently, det X/prime
H=
detXH. This shows explicitly that the group SO( 2, 2)is locally isomorphic to SL( 2,R)⊗SL( 2,R). Although two-
timing theories are bound to be trouble, we could use SO( 2, 2)formally in computing scattering amplitudes, as
we will see in chapter N.3.
Topological quantization of helicity
As promised, let us go back to the observation in the previous section that the matrix XE/((x4)2+/vectorx2)1
2is an
element of SU( 2). Define wA≡xA/((x4)2+/vectorx2)1
2forA=1, 2, 3, 4. An arbitrary element of SU( 2)can be written
asU=w4I+i/vectorw./vectorσ, with det U=1=(w4)2+/vectorw2. The 4-dimensional unit vector w=(w4,/vectorw)traces out the
3-sphere S3, the surface of the 4-ball B4living in 4-dimensional Euclidean space. Thus the group manifold of
SU( 2)isS3.
Next, recall that SU( 2)double covers the rotation group SO( 3), or in plain talk, two elements Uand−U of
SU( 2)corresponds to the same rotation. Thus the group manifold of SO( 3)isS3/Z2, that is, the 3-sphere with
antipodal points identified.
Appendix B. Brief Review of Group Theory | 533
Consider closed paths in SO( 3). Starting at some point PonS3, wander off a bit and come back to P. The
path you traced can evidently be continuously shrunk to a point. But suppose you go off to the other side ofthe world and arrive at −P, the antipodal point of P. You also trace a closed path in SO( 3)since Pand−P
correspond to the same element of SO( 3), but this closed path obviously cannot be shrunk to a point. On the
other hand, if after arriving at −P you keep going and eventually return to P, then the entire path you traced
can be continuously shrunk to a point. Using the language of homotopy groups introduced in chapter V .7, wesay that /Pi1
1(SO(3 ))=Z2: there are two topologically inequivalent classes of paths in the 3-dimensional rotation
group.
Now we can go back and tie up a loose end in chapter III.4. Back in school you learned that the nonlinear
algebraic structure of the Lie algebra [ Ji,Jj]=i/epsilon1ijkJkenforces quantization of angular momentum. But the little
group for a massless particle is merely O(2). In the “rich man’s approach” to gauge invariance, how do we get
the helicity of the photon and the graviton quantized?
The answer is that we invoke topological, rather than algebraic, quantization. A rotation through 4 πis
represented by ei4πhon the helicity hstate of the massless particle, but the path traced out by this rotation
can be continuously shrunk to a point. Hence, we must have ei4πh=1 and h=0,±1
2,±1 ,... .
Appendix C Feynman Rules
Here we gather the Feynman rules given in various chapters.
Draw all possible diagrams. Label each line with a momentum. If applicable, also label each line with an
incoming and an outgoing Lorentz index (for a line describing a vector field), with an incoming and an outgoinginternal index (for a line describing a field transforming under an internal symmetry), so on and so forth.Momentum is conserved at each vertex. Momenta associated with internal lines are to be integrated over withthe measure/integraltext
[d
4p/(2π)4]. A factor of (−1)is to be associated with each closed fermion loop. External lines
are to be amputated. For an incoming fermion line write u(p,s)and for an outgoing fermion line ¯u(p/prime,s/prime).
For an incoming antifermion, write ¯v(p,s), and for an outgoing antifermion, v(p/prime,s/prime). If there are symmetry
transformations leaving the diagram invariant, then we have to worry about the infamous symmetry factors.Since I don’t trust the compilations in various textbooks I work out the symmetry factors from scratch, and thatis what I advise you to do.
Scalar field interacting with Dirac field
L=¯ψ(iγμ∂μ−m)ψ+1
2[(∂ϕ)2−μ2ϕ2]−λ
4!ϕ4+fϕ¯ψψ (1)
Scalar propagator:
ki
k2−μ2+iε(2)
Appendix C. Feynman Rules | 535
Scalar vertex:
−iλ (3)
Fermion propagator:
p i
/negationslashp−m+iε=i/negationslashp+m
p2−m2+iε(4)
Scalar fermion vertex:
if (5)
Initial external fermion:
u(p,s) (6)
Final external fermion:
¯u(p,s) (7)
Initial external antifermion:
¯v(p,s) (8)
Final external antifermion:
v(p,s) (9)
Vector field interacting with Dirac field
L=¯ψ(iγμ(∂μ−ieAμ)−m)ψ−1
4FμνFμν−1
2μ2AμAμ(10)
Vector boson propagator:
k i
k2−μ2/parenleftbiggkμkν
μ2−gμν/parenrightbigg
(11)
Photon propagator (with ξan arbitrary gauge parameter):
k i
k2/bracketleftbigg
(1−ξ)kμkν
k2−gμν/bracketrightbigg
(12)
536 | Appendix C. Feynman Rules
Vector boson fermion vertex:
μ
ieγμ(13)
Initial external vector boson:
εμ(k) (14)
Final external vector boson:
εμ(k)∗(15)
Nonabelian gauge theory
Gauge boson propagator:
k i
k2/bracketleftbigg
(1−ξ)kμkν
k2−gμν/bracketrightbigg
δab (16)
Ghost propagator:
k i
k2δab (17)
Cubic interaction between the gauge bosons:
a,μ
c,λ b,νk1
k3k2gfabc[gμν(k1−k2)λ+gνλ(k2−k3)μ+gλμ(k3−k1)ν] (18)
Quartic interaction between the gauge bosons:
a,μb,ν
d,ρc,λ−ig2[fabefcde(gμλgνρ−gμρgνλ)
+fadefcbe(gμλgνρ−gμνgρλ)
+facefbde(gμνgλρ−gμρgνλ)](19)
Appendix C. Feynman Rules | 537
Gauge boson coupling to the ghost field:
c,μ
abpgfabcpμ(20)
Cross sections and decay rates
Given the Feynman amplitude Mfor a process p1+p2→k1+k2+...+knthe differential cross section is
given by
dσ=1
|/vectorv1−/vectorv2|E(p1)E(p2)d3k1
(2π)3E(k1)...d3kn
(2π)3E(kn)(2π)4δ(4)(p1+p2−n/summationdisplay
i=1ki)|M|2(21)
Here/vectorv1and/vectorv2denote the velocities of the incoming particles. The energy factor E(p)=2/radicalbig
/vectorp2+m2for bosons
andE(p)=/radicalbig
/vectorp2+m2/mfor fermions come from the different normalization of the creation and annhilation
operators in chapters I.8 and II.2.
For a decay of a particle of mass Mthe differential decay rate in its rest frame is given by
d/Gamma1=1
2Md3k1
(2π)3E(k1)...d3kn
(2π)3E(kn)(2π)4δ(4)(P−n/summationdisplay
i=1ki)|M|2(22)
Appendix D Various Identities and Feynman Integrals
Gamma matrices
Identities for the trace of a product of an even number of gamma matrices:
trγμγν=4ημν(1)
trγμγνγλγσ=4(ημνηλσ−ημληνσ+ημσηνλ) (2)
We define the totally antisymmetric symbol εμνλσbyε0123=+ 1 (note ε0123=− 1). Then with our definition
γ5≡iγ0γ1γ2γ3, we have
trγ5γμγνγλγσ=− 4iεμνλσ(3)
Identities that follow from the basic Clifford identity:
γμ/negationslashpγμ=− 2/negationslashp (4)
γμ/negationslashp/negationslashqγμ=4p.q (5)
γμ/negationslashp/negationslashq/negationslashrγμ=− 2/negationslashr/negationslashq/negationslashp (6)
I leave it to you to derive these identities. For example, to obtain (4) keep moving γμto the right in the expression
γμ/negationslashpγμ=(2pμ− /negationslashpγμ)γμ=2/negationslashp−4/negationslashp=− 2/negationslashp.
Evaluating Feynman diagrams
Over the years, a number of tricks and identities have been developed for evaluating the integrals associated with
Feynman diagrams.
Let us evaluate
I=/integraldisplayd4k
(2π)41
(k2−m2+iε)3=/integraldisplayd3k
(2π)3/integraldisplaydk0
2π1
[k2
0−(/vectork2+m2)+iε]3
Focus on the k0integral. Draw where the poles are in the complex k0-plane and you will see that the integration
contour can be rotated anticlockwise so that [we denote the integrand by f( k 0)]
/integraldisplay+∞
−∞dk0f( k 0)=/integraldisplay+i∞
−i∞dk0f( k 0)=i/integraldisplay+∞
−∞dk4f( ik4) (7)
Appendix D. Identities and Feynman Integrals | 539
where in the last step we define k0=ik4(corresponding to the Wick rotation mentioned in chapters I.2 and V .2.)
Thus,
I=i(−1)3/integraldisplayd4
Ek
(2π)41
(k2
E+m2)3
where d4
Ekis the integration element in Euclidean 4-dimensional space and k2
E≡k2
4+/vectork2the square of a Euclidean
4-vector. The infinitesimal εcan now be set equal to zero. We can integrate immediately over the three angles
since the integrand does not depend on them. You can look up the angular element in Euclidean space in a book,but we will use a neat trick instead.
I will do the more general d-dimensional integral H=/integraltext
d
dkF(k2), where k2=k2
1+k2
2+...+k2
dandFcan
be any function as long as the integral converges. (I now drop the subscript E; the context makes clear that we
are in Euclidean space.) We can of course set dequal to 4 at the end. The result for arbitrary dwill be useful to
us in regularizing dimensionally (chapter III.1).
We imagine integrating over the (d−1)angular variables to obtain H=C(d)/integraltext∞
0dk
kd−1F(k2). To determine C(d) we will do the integral J=/integraltext
ddke−1
2k2in two different ways. Using (I.2.8) we
haveJ=(√
2π)d. Alternatively,
J=C(d)/integraldisplay∞
0dk kd−1e−1
2k2=C(d) 2d
2−1/integraldisplay∞
0dx xd
2−1e−x=C(d) 2d
2−1/Gamma1(d
2)
where we changed integration variables and recognized the integral representation of the gamma function
/Gamma1(z+1)≡/integraltext∞
0dx xze−x. (Recall that upon integration by parts we obtain /Gamma1(z+1)=z/Gamma1(z), so that /Gamma1(n)=(n−1)!
fornan integer.) Therefore C(d)=2πd/2//Gamma1(d/2 )and
/integraldisplay
ddkF(k2)=2πd/2
/Gamma1(d/ 2)/integraldisplay∞
0dk kd−1F(k2) (8)
Setting d=1 in (8) we determine /Gamma1(1
2)=π1
2, and setting F(k2)=δ(k−1)we see that the area of the (d−1)-
dimensional sphere is equal to C(d) , thus recovering various results you learned in school about circles and
spheres: C(2)=2πandC(3)=4π.
The new result you need as a budding field theorist is for d=4:
/integraldisplay
d4kF(k2)=π2/integraldisplay∞
0dk2k2F(k2) (9)
So finally we have
I=−i
16π2/integraldisplay∞
0dk2k2 1
(k2+m2)3=−i
16π21
2m2(10)
We have derived the basic formula for doing Feynman integrals:
/integraldisplayd4k
(2π)41
(k2−m2+iε)3=−i
32π2m2(11)
(With the telltale iεwe have evidently moved back to Minkowski space.) As an exercise you can go through the
same steps to find
/integraldisplay/Lambda1d4k
(2π)41
(k2−m2+iε)2=i
16π2/bracketleftbigg
log/parenleftbigg/Lambda12
m2/parenrightbigg
−1+.../bracketrightbigg
(12)
Here a cutoff is needed, which we introduce by setting the upper limit in the integral over k2in the analog of
(10) to /Lambda12. As a check, differentiate (12) with respect to m2to recover (11). As another exercise show that
/integraldisplay/Lambda1d4k
(2π)4k2
(k2−m2+iε)2=−i
16π2/bracketleftbigg
/Lambda12−2m2log/parenleftbigg/Lambda12
m2/parenrightbigg
+m2+.../bracketrightbigg
(13)
In (12) and (13) (...)denote terms that vanish for /Lambda12/greatermuchm2. In some texts, the (−1)in (12) is dropped by absorbing
it into /Lambda12. But then we have to be careful to adjust (13) accordingly if it appears in the same calculation.
540 | Appendix D. Identities and Feynman Integrals
A useful identity in combining denominators is
1
x1x2...xn=(n−1)!/integraldisplay1
0/integraldisplay1
0.../integraldisplay1
0dα1dα2...dαn
δ⎛
⎝1−n/summationdisplay
jαj⎞
⎠1
(α1x1+α2x2+...+αnxn)n(14)
Forn=2,
1
xy=/integraldisplay1
0dα1
[αx+(1−α)y]2(15)
and for n=3,
1
xyz=2/integraldisplay1
0/integraldisplay1
0/integraldisplay1
0dαdβdγδ(α +β+γ−1)1
(αx+βy+γz)3(16)
=2/integraldisplay/integraldisplay
triangledαdβ1
[z+α(x−z)+β(y−z)]3
where the integration region is the triangle in the α-βplane bounded by 0 ≤β≤1−αand 0 ≤α≤1.
Appendix E Dotted and Undotted Indices and the Majorana Spinor
We develop the dotted and undotted notation introduced in chapter II.3 for further use in discussing supersym-
metry in chapter VIII.4 and in part N. In essence, the appearance of undotted and dotted indices can be traced
back to the fact that the algebra of the Lorentz group SO( 3, 1), with the generators /vectorJ+i/vectorKand/vectorJ−i/vectorK, breaks
up into two pieces, each isomorphic to the algebra of SU( 2). The absence or presence of the dot allows us to keep
track of which SU( 2)we are talking about.
Here I will use extensively results from chapter II.3 and from the exercises (do them!) there without bothering
to write them down again here.
In the Weyl basis of chapter II.1
γμ=/parenleftBigg0σμ
¯σμ0/parenrightBigg
(1)
where σμ=(I,/vectorσ)and¯σμ=(I,−/vectorσ). Knowing that γμacts on
/Psi1=/parenleftBiggψα
¯χ˙α/parenrightBigg
we see that σμand¯σμcarry indices as follows:
(σμ)α˙αand(¯σμ)˙αα(2)
This is consistent with what you know: the Lorentz vector transforms like (1
2,1
2)and thus straddles the two
SU( 2)’s. The matrices σμand¯σμmix dotted and undotted indices. We will make good use of this observation
later.
Let us check that the Lorentz transformation property of the Dirac spinor /Psi1is consistent with what was
discussed in chapter II.1. There we learned that /Psi1→e−i
4ωμν/Sigma1μν/Psi1, where /Sigma1μν≡i
2[γμ,γν]. (We want to use the
symbol σμνfor some other quantity, hence the change of notation.) Using (1) we obtain
/Sigma1μν=2i/parenleftBiggσμν0
0¯σμν/parenrightBigg
where σμν≡1
4(σμ¯σν−σν¯σμ)and¯σμν≡1
4(¯σμσν−¯σνσμ). From (2) we see that these two matrices carry indices
as follows:
(σμν)β
αand(¯σμν)˙α
˙β(3)
542 | Appendix E. Indices and the Majorana Spinor
Again, this reflects the fact that the antisymmetric tensor (such as the electromagnetic field Fμν)transforms like
(1, 0)+(0, 1).
The matrices σμνand¯σμνmay seem alien, but recall that they are manufactured out of the familiar Pauli
matrices and so they are simply Pauli matrices (what else could they be?) themselves. In particular,
σ0i=−¯σ0i=−1
2σiand σij=¯σij=−i
2εijkσk
Note that these relations are consistent with (σμν)†=−(¯σμν), which in turn follows from (/Sigma1μν)†=γ0/Sigma1μνγ0.
Mother Nature is kind to the students of quantum field theory. The relativistic spinor /Psi1breaks up into two
2-component spinors acted on by the Pauli matrices. What you learned in nonrelativistic quantum mechanicscontinues to be relevant here.
Thus under an infinitesimal Lorentz transformation
ψα→/parenleftBig
I+1
2ωμνσμν/parenrightBigβ
αψβ (4)
and
¯χ˙α→/parenleftBig
I+1
2ωμν¯σμν/parenrightBig˙α
˙β¯χ˙β(5)
You should check that it all works out according to plan. Everything is consistent with what we learned in chapter
II.3, in particular, that boosts act oppositely on (1
2,0)and(0,1
2), but rotations act the same.
Thus far, on the spinor fields ψαand¯χ˙α, the dotted indices always live upstairs and the undotted indices
downstairs. What would get them to change floors? Charge conjugation.
Recall from chapter II.1 that the charge conjugated field is defined by /Psi1c≡C¯/Psi1T[where Tdenotes transpose,
¯/Psi1means /Psi1†γ0, andC−1γμC=−(γμ)T.] In the Weyl basis, we can choose
C=ζγ0γ2=ζ/parenleftBigg−σ20
0σ2/parenrightBigg
(6)
The condition (/Psi1c)c=/Psi1implies |ζ|=1. We choose ζ=−i. Explicitly,
/Psi1c=/parenleftBiggiσ2¯χ∗
−iσ2ψ∗/parenrightBigg
We now introduce some notation, the wisdom of which will soon become clear. Given ψαand¯χ˙α, define
¯ψ˙α≡(ψα)∗andχα≡(¯χ˙α)∗(7)
Weird, complex conjugation puts on a dot and a bar.
We raise and lower undotted indices as follows: ψα=εαβψβandψβ=εβγψγwhich implies that εαβεβγ=δγ
α.
Thus, if we choose
εαβ=/parenleftBigg01
−10/parenrightBigg
=(iσ 2)αβ
then
εβγ=/parenleftBigg0−1
10/parenrightBigg
=(−iσ 2)βγ
We are forced to define ε12=+ 1 and ε12=− 1 to have opposite signs, a fact to keep in mind.
You should realize by now that what we are doing can again be traced back to that peculiar fact about Pauli
matrices (appendix B):
(iσ 2)σ∗
i(−iσ 2)=−σi (8)
or equivalently
σ2σT
iσ2=−σi (9)
Appendix E. Indices and the Majorana Spinor | 543
an identity in one guise or another familiar from quantum mechanics. We have used it again and again, in
appendix B and in the text (for example, in connection with Majorana masses and with the Higgs field). From(8) we have (iσ
2)σμ∗(−iσ 2)=¯σμand hence
(iσ 2)(σμν)∗(−iσ 2)=¯σμν(10)
Analogously, we raise and lower dotted indices as follows: ¯ψ˙α=ε˙α˙β¯ψ˙βand¯ψ˙β=ε˙β˙γ¯ψ˙γ. Referring to (7) we
see that ε˙α˙βis numerically the same as εαβ, andε˙β˙γis numerically the same as εβγ.
You now see the rationale of these apparently capricious choices: we can now write
/Psi1c=/parenleftBiggχα
¯ψ˙α/parenrightBigg
(11)
Referring to
/Psi1=/parenleftBiggψα
¯χ˙α/parenrightBigg
(12)
we see that the point of the notation is that ψαandχαtransform in the same way and are the same kind of
creature (and similarly for ¯χ˙αand¯ψ˙α.)
We now come to the all-important concept of a Majorana spinor. Ettore Majorana, a brilliant physicist,
mysteriously disappeared early in his career. Fermi supposedly described Majorana as “a towering giant withoutany common sense.”
1
Given a Dirac spinor /Psi1,i f/Psi1=/Psi1c, then /Psi1is said to be a Majorana spinor.
Comparing (12) and (11), we see that a Majorana spinor has the form
/Psi1M=/parenleftBiggψα
¯ψ˙α/parenrightBigg
(13)
An obvious remark but a handy mnemonic: Given a Weyl spinor ψαwe can construct a Majorana spinor, and
given two Weyl spinors we can construct a Dirac spinor: one Weyl equals one Majorana, and two Weyls equalone Dirac.
Incidentally, another way of seeing that complex conjugation puts on a dot is that (see chapter II.3) conjugation
interchanges /vectorJ+i/vectorKand/vectorJ−i/vectorK.
The point to remember is simply that given a spinor λ
α, then λ˙αtransforms like (λα)∗. You should verify this,
keeping in mind (10).
The utility of the notation is similar to that of the covariant and contravariant (or upper and lower) indices in
special and general relativity. We always contract an upper index with a lower index. Here we have the additionalrule that an undotted upper index can only be contracted with an undotted lower index, but never with a dottedlower index, (obviously, since they belong to different algebras.) It is easy to verify these rules. For example, letus show that η
αψαis invariant. Using (4) we proceed with laboriously careful pedagogy:
ηα→η/primeα=εαβη/prime
β=εαβ(e1
2ωσ)γ
βηγ=εαβ(e1
2ωσ)γ
βεγρηρ=(e−1
2ωσT)α
ρηρ(14)
where we used once again the identity (9). Then ηαψα→η(e−1
2ωσT)T(e1
2ωσ)ψ=ηψ, which is indeed an invariant.
In special and general relativity we raise and lower indices with the metric, which is of course symmetric.
Here we raise and lower indices with the antisymmetric εsymbol and as a result signs pop up here and there.
For example, ηαψα=εαβηβψα=ηβ(−εβα)ψα=−ηβψβ. Contrast this with the scalar product of two vectors
vμwμ=vμwμ. If we want to suppress indices and write ηψ, we must decide once and for all what that means.
The standard convention is to define
ηψ≡ηαψα (15)
and not ηβψβ. This rule is sometimes stated by saying that in contracting undotted indices we always go from
the northwest to the southeast, and never from southeast to northwest. As we learned in chapter II.5, spinorfields are to be treated as anticommuting Grassman variables under the path integral, so that −η
βψβ=ψβηβ.
We end up with the nice rule ηψ=ψη.
1M. Gell-Mann, private communication. Incidentally, the name Ettore corresponds to Hector in English.
544 | Appendix E. Indices and the Majorana Spinor
Similarly, we define
¯χ¯ξ≡¯χ˙α¯ξ˙α=¯ξ¯χ (16)
In contracting dotted indices we always go from southwest to northeast. Of course, none of this “Santa Barbara
to Cambridge” convention is needed if the indices are displayed explicitly.
Just as in special and general relativity, where the upper and lower indices are very useful in telling us
whether expressions we write down make sense, the undotted and dotted upper and lower indices allow usto see immediately that ηψandησ
μνψmake sense, but that ησμψdoes not. [Look at (2) and (3) and notice the
kind of indices that appear.] The notation of course just codifies in a convenient way the group theory fact that
(1
2,0)⊗(1
2,0)=(0, 0)⊕(1, 0), namely, that out of two Weyl spinors we can make a scalar and a tensor but not
a vector.
As always, notation should be driven by physics and computational convenience (which is intimately con-
nected to elegance).
To gain familiarity with the dotted and undotted 2-component notation, you should work out some of the
identities in the exercises. These identities are useful when working with supersymmetric field theories.
Exercises
E.1 Show that ησμνψ=−ψσμνηand¯χ¯σμψ=−ψσμ¯χ.
E.2 Show that (θϕ)(¯χ¯ξ)=−1
2(θσμ¯ξ)(¯χ¯σμϕ).
E.3 Show that θαθβ=1
2(θθ)δα
β. [Hint: simply evaluate the two sides for all possible cases.]
Solutions to Selected Exercises
Part I
I.3.1 From the text we have for x0=0,
D(x)=−i/integraldisplayd3k
(2π)32/radicalbig
/vectork2+m2e−i/vectork./vectorx
=−i
2(2π)2/integraldisplay∞
0dkk2
/radicalbig
/vectork2+m2/integraldisplay+1
−1d(cosθ)e−ikr cosθ
=−1
2(2π)2r/integraldisplay∞
0dkk/radicalbig
/vectork2+m2(eikr−e−ikr)=−1
8π2r/integraldisplay∞
−∞dkk/radicalbig
/vectork2+m2eikr
=i
8π2r∂
∂r/integraldisplay∞
−∞dk/radicalbig
/vectork2+m2eikr
The integrand in I≡/integraltext∞
−∞(dk//radicalbig
/vectork2+m2)eikrhas a cut along the imaginary axis going from imtoi∞(and
another cut we don’t care about.) So fold the contour around the cut and change variable to k=i(m+y):
I=2/integraldisplay∞
0dye−(y+m)r 1/radicalbig
(y+m)2−m2
=2/integraldisplay∞
1due−mru 1√
u2−1
=2/integraldisplay∞
0dte−mr cosht.
546 | Solutions to Selected Exercises
At this point you can look in a table and find that this is some Bessel function and read off the large r
behavior, but it is more stylish to press on and descend steeply: we obtain
D(x)=−im
4π2r/integraldisplay∞
0dt(cosht)e−mr cosht
=−im
4π2r/integraldisplay∞
0d(sinht)e−mr cosht
=−im
4π2r/integraldisplay∞
0dse−mr√
s2+1/similarequal−im
4π2r/integraldisplay∞
0dse−mr(1 +1
2s2)
=−im2
4π2(π
2(mr)3)1
2e−mr,
using the Gaussian integral from the appendix of chapter I.2.
I.3.2 We evaluate
D(x)=/integraldisplayd2k
(2π)2eikx
k2−m2+iε
by contours as in the text and obtain
D(x)=−i/integraldisplaydk
(2π)2ωk[e−i(ω kt−kx)θ(x0)+ei(ωkt−kx)θ(−x0)]
Forx0=0, we recognize the integral
D(x)=−i/integraldisplay+∞
−∞dk
(2π)2/radicalbig
/vectork2+m2e−ikx
as a Bessel function from exercise I.3.1:
D(x)=−i
2πK0(m|x|)→−i
2π/radicalbiggπ
2m|x|e−m|x|
with the expected exponential decay for large x.
I.7.2 Expanding and keeping only the desired terms
Z(J)→C/braceleftBigg
1+1
2!/parenleftbigg
−i
4!λ/parenrightbigg2/integraldisplay/integraldisplay
d4w1d4w2/bracketleftbiggδ
iδJ(w1)/bracketrightbigg4/bracketleftbiggδ
iδJ(w2)/bracketrightbigg4
1
6!/bracketleftbigg
−i
2/integraldisplay/integraldisplay
d4xd4yJ(x)D(x −y)J(y)/bracketrightbigg6/bracerightBigg
Just keep on differentiating.
I.7.4 Write k1=(√
k2+m2,0 ,0 ,k) andk2=(√
k2+m2,0 ,0 ,−k) . Then E=2√
k2+m2≥2m. Physically, a
pair of mesons can be produced when E≥2m.
I.8.1 Do the k0integral on the left-hand side of (I.8.14):/integraltext
dk0δ((k0)2−ω2
k)θ(k0)f (k0,/vectork), where ωk≡
+/radicalbig
/vectork+m2. Using (I.2.12) and picking up the positive root because of the step function, we obtain/integraltext∞
0dk0(δ(k0−ωk)/(2k0))f (k0,/vectork)=f( ωk,/vectork)/(2ωk).
To verify the invariance explicitly, boost in the xdirection and drop the subscript on ωk:kx→
sinhφω+coshφkxandω→coshφω+sinhφkx. Then, using ω2=(kx)2+...and hence ωdω=
kxdkx, we have dkx→(sinhφ( kx/ω)+coshφ)dkx. Hence dkx/ω→dkx/ω.
Solutions to Selected Exercises | 547
I.8.2 Clearly, only the terms aa†anda†ainHcontributes to </vectork/prime|H|/vectork> . Extract these two types of terms in
/integraldisplay
dDxϕ(x)2
=/integraldisplay
dDx/integraldisplay/integraldisplaydDq/radicalBig
(2π)D2ωqdDq/prime
/radicalBig
(2π)D2ωq/prime[a(/vectorq)a†(/vectorq/prime)e−i(ω qt−/vectorq./vectorx)ei(ωq/primet−/vectorq/prime./vectorx)+h.c.]
=/integraldisplaydDq
2ωq[a(/vectorq)a†(/vectorq)+a†(/vectorq)a(/vectorq)]
and so His for our purposes effectively equal to/integraltext
dDqωq
2[a(/vectorq)a†(/vectorq)+a†(/vectorq)a(/vectorq)], which upon using
the commutation relation is equal to/integraltext
dDqωq
2[δ(D)(/vector0)+2a†(/vectorq)a(/vectorq)]. We recognize the first term as the
vacuum calculated in the text. Note that the definition of the delta function (2π)Dδ(D)(/vectork)=/integraltext
dDxei/vectork./vectorx
implies δ(D)(/vector0)=[1/(2π)D]/integraltext
dDx=V/(2π)D. Thus, subtracting off the vacuum energy, we have H
effectively equal to/integraltext
dDqωqa†(/vectorq)a(/vectorq), which just says that a mode of momentum /vectorqcarries energy ωq.
In particular, using the commutation relation twice we have </vectork/prime|H|/vectork>=δ(D)(/vectork/prime−/vectork)ωk. The energy of
a particle of momentum /vectorkisωkrelative to the vacuum.
I.8.4 Q=/integraltext
dDxJ0(x)=/integraltext
dDx(ϕ†i∂0ϕ−i(∂0ϕ†)ϕ). Focus on the first term:
/integraldisplay
dDx/integraldisplay/integraldisplaydDk/prime
/radicalbig
(2π)D2ωk/primedDk/radicalbig
(2π)D2ωk
[a†(/vectork/prime)ei(ωk/primet−/vectork/prime./vectorx)+b(/vectork/prime)e−i(ωk/primet−/vectork/prime./vectorx)]ωk[a(/vectork)e−i(ω kt−/vectork./vectorx)−b†(/vectork)ei(ωkt−/vectork./vectorx)]
Note that i∂0brings down a factor of ωkand produces a relative sign between aandb†. As in exercise I.8.2
the integral over xproduces a delta function that collapses the two kintegrals into one, giving
/integraldisplay
dDk1
2(a†(/vectork)a(/vectork)−b(/vectork)b†(/vectork)−a†(−/vectork)b†(/vectork)e2iωkt+b(−/vectork)a(/vectork)e−2iωkt)
The second term in −i(∂ 0ϕ†)ϕinJ0(x)is just the hermitean conjugate of the first term ϕ†i∂0ϕ. Thus,
adding the hermitean conjugate of what we have just obtained, we find
Q=/integraldisplay
dDk[a†(/vectork)a(/vectork)−b(/vectork)b†(/vectork)]
=/integraldisplay
dDk[a†(/vectork)a(/vectork)−b†(/vectork)b(/vectork)]+δ(D)(/vector0)/integraldisplay
dDk
The infinite additive constant is to be subtracted out much like the vacuum energy. In some texts a normal
ordering operation, denoted by a pair of colons, is defined as follows: If you see : (...): you are instructed
to move all the creation operators in the expression (...)to the left of the annihilation operators. In other
words, by fiat : b(/vectork)b†(/vectork):≡b†(/vectork)b(/vectork). The current is then defined by Jμ(x)≡:(ϕ†i∂μϕ−i(∂μϕ†)ϕ) :.
Since the normal ordered current differs from the naively defined current by a c-number the most crucial
property of the current, namely current conservation ∂μJμ=0, is not affected. This is of course just a
formal way of saying that the value of the charge in the vacuum state is to be subtracted. In any case, the
result Q=/integraltext
dDk[a†(/vectork)a(/vectork)−b†(/vectork)b(/vectork)] shows that aandbannihilate positive and negative charges,
respectively.
I.10.2 We have (repeated indices summed)
Raa/primeRbb/primeiDa/primeb/prime(x)=/integraldisplay
DϕR aa/primeϕa/prime(x)R bb/primeϕb/prime(0)eiS
But we can change the integration variable from ϕtoRϕ. Since the action Sand the measure Dϕ are
both invariant under SO(N) rotations, this is equal to/integraltext
Dϕϕ a(x)ϕb(0)eiS=iDab(x). Thus, we obtain
Dab=Raa/primeRbb/primeDa/primeb/prime. The properties of the rotation group are such that the only solution of this equation
isDabproportional to δab.
548 | Solutions to Selected Exercises
I.10.3 The field ϕtransforms as a symmetric traceless tensor (see appendix B) under SO( 3), that is, with all
indices displayed, ϕab→Raa/primeRbb/primeϕa/primeb/prime=Raa/primeϕa/primeb/primeRT
b/primeb=(RϕRT)ab. As suggested in the hint, writing ϕ
as a 3 by 3 symmetric traceless matrix we have ϕ→RϕRTand thus the invariants are (up to quartic
order in ϕ)tr(∂μϕ)2,t rϕ2,t rϕ4, and (tr ϕ2)2. Remarkably, you can prove that tr ϕ4and (tr ϕ2)2actually
amount to only one invariant by diagonalizing
ϕ=⎛
⎜⎜⎝α 00
0β 0
00 −(α+β)⎞
⎟⎟⎠
You can see by computation tr ϕ4and (tr ϕ2)2are both proportional to [ α2+β2+(α+β)2]2. Thus, if
we restrict ourselves to quartic terms the Lagrangian L=1
2tr(∂μϕ)2−1
2m2trϕ2−λ(trϕ2)2actually has
anSO( 5)symmetry (since ϕhas 5 components.) This is an example of what is known as “accidental
symmetry.” Convince yourself that this holds only to quartic order in ϕ.
I.11.2 Varying gμρgρλ=δμ
λwe have ( δgμρ)gρλ=−gμρ(δgρλ), which upon multiplication by gλνbecomes
δgμν=−gμρ(δgρλ)gλν. You may recognize this as just the statement δM−1=−M−1(δM)M−1for a
matrix M. To evaluate δgwe use the important identity det M=eTr log M, which you can prove easily
by diagonalizing Mwith a similarity transformation. The left hand side is equal to the product of the
eigenvalues of M, while the right hand side is equal to the exponential of the sum of logarithms of the
eigenvalues. {You can define the logarithm of a matrix by expanding log[ I+(M−I)] in a power series
in(M−I).}
Thus, δdetM=(detM)trM−1δM and so δg=ggνμδgμν. We are now ready to vary
S=/integraldisplay
d4x√−g1
2(gμν∂μϕ∂νϕ−m2ϕ2)≡/integraldisplay
d4x√−gL
Plugging in, we have
δS=/integraldisplay
d4x√−g[1
2gνμδgμνL−gμρ(δgρλ)gλν1
2∂μϕ∂νϕ]
Thus,
Tμν=−2√−gδS
δgμν=gμρgνλ∂ρϕ∂λϕ−gμνL
In the flat spacetime limit
T00=(∂0ϕ)2−L=1
2((∂0ϕ)2+(/vector∇ϕ)2+m2ϕ2)
precisely the energy density as promised.
I.11.3 Using the expression for Tμνfrom the preceding exercise, we have
Pi=/integraldisplay
d3xT0i=−/integraltext
d3x∂0ϕ∂iϕ
and
[Pi,ϕ(x) ]=−/integraldisplay
d3y[∂0ϕ(y) ,ϕ(x) ]∂iϕ(y)=i∂iϕ(x)
Thus, combined with the fact that P0=H, we have [Pμ,ϕ(x) ]=−i∂μϕ(x) , which just reflects the fact
thatPμandxνare conjugate variables.
I.11.4 Evaluating Tμν=−FμλFλ
ν−ημνL, we have
Tij=−FiλFλ
j+1
2δij(/vectorE2−/vectorB2)=−EiEj+FikFjk+1
2δij(/vectorE2−/vectorB2)
Solutions to Selected Exercises | 549
SinceFikFjk=εikmεjknBmBn=δij/vectorB2−BiBj, we obtain the announced result. Note that δijTij=1
2(/vectorE2+
/vectorB2)=T00and hence T=0.
Part II
II.1.1 Continuing the hint, we have
δ(¯ψγμγ5ψ)=¯ψi
4ωλρ[σλρ,γμγ5]ψ=¯ψi
4ωλρ[σλρ,γμ]γ5ψ
since γ5anticommutes with gamma matrices and hence commutes with the product of two gamma
matrices. Inserting [ σλρ,γμ] as given in the text we have δ(¯ψγμγ5ψ)=ωμ
λ¯ψγλγ5ψ, which is precisely
how a vector transforms. Under parity ¯ψγμγ5ψ→¯ψγ0γμγ5γ0ψwhich equals ¯ψγ5γ0ψ=−¯ψγ0γ5ψ
forμ=0, and ¯ψγ0γiγ5γ0ψ=¯ψγiγ5ψforμ=i. The time component flips sign while the spatial
components do not. Thus, the behavior under parity is opposite to that of a normal vector: ¯ψγμγ5ψ
is an axial vector. The other cases proceed similarly.
II.1.2 FromψL=1
2(1−γ5)ψandψR=1
2(1+γ5)ψ, we find ¯ψL=ψ†
Lγ0=ψ†1
2(1−γ5)γ0=¯ψ1
2(1+γ5)and
¯ψR=¯ψ1
2(1−γ5). We then just use the properties of PLandPRrepeatedly. For example, ¯ψLψR=¯ψ1
2(1+
γ5)ψand¯ψRψL=¯ψ1
2(1−γ5)ψ, or equivalently, ¯ψψ=¯ψLψR+¯ψRψLand¯ψγ5ψ=¯ψLψR−¯ψRψL.A s
another example, ¯ψLγμψL=¯ψ1
2(1+γ5)γμ1
2(1−γ5)ψ=¯ψγμ1
2(1−γ5)ψand¯ψRγμψR=¯ψγμ1
2(1+
γ5)ψ. Note that various combinations vanish, for example, ¯ψLψL=0,¯ψLγμψR=0, and so on. Complete
the exercise.
II.1.3–4 In the appropriate basis the Dirac equation becomes
/parenleftBiggE−mp σ3
−pσ 3−E−m/parenrightBigg/parenleftBiggφ
χ/parenrightBigg
=0
that is, (E−m)φ+pσ3χ=0 and −pσ 3φ−(E+m)χ=0. The second equation informs us that χ=
−[p/(E+m)]σ3φ. For a slow electron χ/similarequal−(p/2m)σ 3φ, so that χis smaller than φby the factor p/2m.
The first equation then reduces to (E−m−p2/2m)φ=0, which just reminds us of the relation between
energy and momentum in the nonrelativistic limit.
II.1.5 In the Weyl basis the Dirac equation for a relativistic electron moving along the 3-axis E(γ0−γ3)ψ=0
becomes
/parenleftBigg0 I−σ3
I+σ30/parenrightBigg/parenleftBiggψL
ψR/parenrightBigg
=0
Since
σ12≡i
2[γ1,γ2]=−i
2[σ1,σ2]⊗I=σ3⊗I=/parenleftBiggσ30
0σ3/parenrightBigg
under a rotation around the 3-axis, ψL→e−(i/4 )ωσ3ψL=e+(i/4 )ωψLwhile ψR→e−(i/4 )ωσ3ψR=
e−(i/4 )ωψR. Indeed, the left and right handed fields rotate in opposite directions.
II.1.6 In the Weyl basis, the Dirac equation γ.pu=0 becomes σμpμη=0, and ¯σμpμχ=0, with
u=/parenleftBiggχ
η/parenrightBigg
The solutions are
η=/parenleftBiggp1−ip2
p0−p3/parenrightBigg
and η=/parenleftBiggp0+p3
p1+ip2/parenrightBigg
550 | Solutions to Selected Exercises
for the two possible helicities. The corresponding solutions for χmay be obtained by /vectorp↔−/vectorp. We have
¯uu=p.p=0. For a particle moving in the +3 direction, η=0 and
χ=2E/parenleftBigg0
1/parenrightBigg
and
η=2E/parenleftBigg1
0/parenrightBigg
andχ=0. The Lorentz vector ¯uγμu=(2E)2(1, 0, 0, 1 ). (What other direction could it point in?) For
a particle moving in the −3 direction, ηandχexchange roles. This exercise shows explicitly that for
massless particles we can use 2-component spinors. (What happens if parity is broken?)
II.1.8 In either the Dirac or the Weyl basis, (ψc)c=γ2(γ2ψ∗)∗=ψ.
II.1.9 It is easiest to work in either the Dirac or the Weyl basis. Let ψbe left handed, that is (1+γ5)ψ=0.
Then(1−γ5)ψc=(1−γ5)γ2ψ∗=γ2(1+γ5)ψ∗=0 since γ5is real.
II.1.10 ψCψ →ψe−i
4ωλρ(σλρ)TCe−i
4ωμνσμνψ=ψCψ since(σλρ)TC=−Cσλρ.
II.1.12 Under parity or reflection in a mirror, x1→x1andx2→−x2. Choose γ0=σ3,γ0γ1=σ1, andγ0γ2=
σ2. Multiply the Dirac equation (iγμ∂μ−m)ψ=0b yγ0and write [ i(∂0+γ0γi∂i)−γ0m]ψ=0. Then
multiplying by σ1reverses the sign of the ∂2term, but also the mass term. I leave it to you to discuss
time reversal.
II.2.1 Apply Noether for the transformation
ψ→eiθψ=(1+iθ)ψ
Then
δL
δ(∂μψ)δψ+δL
δ(∂μ¯ψ)δ¯ψ=¯ψiγμ(iθψ)
Note that formally Ldoes not depend on ∂μ¯ψ. Thus, up to overall factors we can choose Jμ=¯ψγμψ
with the corresponding charge Q=/integraltext
d3x¯ψγ0ψinto which we plug (II.2.10)
ψ(x)=/integraldisplayd3p
(2π)3/2(Ep/m)1/2/summationdisplay
s[b(p,s)u(p ,s)e−ipx+d†(p,s)v(p ,s)eipx]
At this point, the calculation pretty much parallels what you did in exercise I.8.4. The integration/integraltext
d3x
over space produces a delta function that sets the momentum variables in ψand in ¯ψequal to one
another. The new feature here is that we encounter objects such as ¯uγ0u. Invoking Lorentz invariance and
referring to the rest frame form of uandvwe have ¯u(p,s)γμu(p,s/prime)=δss/primepμ/m,¯u(p,s)γμv(p,s/prime)=0,
and so on. We obtain
Q=/integraldisplayd3p
(2π)3(Ep/m)/summationdisplay
s[b†(p,s)b(p ,s)+d(p ,s)d†(p,s)]
As in exercise I.8.4 we have to move the creation operator d†to the left of the annihilation operator d
and subtract off an infinite constant. Thus, finally
Q=/integraldisplayd3p
(2π)3(Ep/m)/summationdisplay
s[b†(p,s)b(p ,s)−d†(p,s)d(p ,s)]
showing clearly that bannihilates a negative charge and da positive charge.
Solutions to Selected Exercises | 551
To calculate [ Q,ψ(0)]=/integraltext
d3x[¯ψ(x)γ0ψ(x) ,ψ(0)] we use the identity [ AB ,C]=A{B ,C}−{A,C}B
and the canonical anticommutation relation (II.2.4). We find [ Q,ψ(0)]=−ψ(0), thus showing that b
andd†must carry the same charge.
II.3.4 The desired equations are γμ/Psi1αμ=0 (this takes out 4 components since αtakes on 4 values) and
(/negationslashp−m)β
α/Psi1βμ=0 (for each μthis takes out 2 components and so altogether 4 ×2=8 components.)
Thus, 16 −4−8=4 components as desired. Another way of saying this is that γμ/Psi1αμis a Dirac spinor
and hence the spin1
2part of the vector-spinor /Psi1αμ.
II.6.4 It is good practice to be as symmetrical as one can in calculations. So define p3≡−P1andp4≡−P2and
add the 6 (not 3) combinations appearing in the definitions of s,t, andu, thus obtaining
2(s+t+u)=(p1+p2)2+(p3+p4)2+(p3+p1)2+(p4+p2)2+(p4+p1)2+(p3+p2)2
=34/summationdisplay
i=1m2
i+2(p1.p2+p3.p4+p1.p3+p2.p4+p1.p4+p2.p3)
The second group of terms on the right-hand side collect into (/summationtext4
i=1pi)2−/summationtext4
i=1m2
i. (Obviously, we
have for convenience changed notation slightly, setting m3=M1andm4=M2.)
II.6.5 Referring to (C.11) we see that in dσthe factor
1
|/vectorv1−/vectorv2|E(p1)E(p2)1
(2π)3E(k1)...1
(2π)3E(kn)(2π)4
reduces to1
2(m/E)4[1/(2π)2]. Integrating the factor d3P1d3P2δ(4)(p1+p2−P1−P2)over/vectorP2we knock
off 3 of the delta functions, leaving us with d/Omega1dP1P2
1δ(2E−E1), and so the integral over P1givesd/Omega11
2E2.
Finally, the factor containing the “real physics” is1
2/summationtext
s/summationtext
S|M|2=(e4/4m4)f (θ) . Multiplying the three
factors together and dividing by d/Omega1, we obtain dσ/d/Omega1 =(1
2)5[e4/(2π)2](1/E2)f (θ), as given in the text.
Note that mcancels out as expected. We should be able to take the limit m→0 compared to the energies
without the cross section either blowing up or vanishing.
II.6.7/Gamma1=|M|2
2M/integraldisplayd3k
(2π)32ωd3k/prime
(2π)32ω/prime(2π)4δ4(k+k/prime−q)
Knock off the /vectork/primeintegral and do the angular part of the /vectorkintegral to obtain
/Gamma1=|M|2
8πM/integraldisplaydkk2
ωω/primeδ(/radicalbig
k2+m2+/radicalbig
k/prime2+m2−M)
Using (I.2.12), we evaluate the integral as (k2/ωω/prime)(1/(k
ω+k
ω/prime))=k/M . Solving√
k2+m2+√
k/prime2+m2=
Mforkwe obtain the stated result.
Part III
III.1.2 The amplitude should become nonanalytic when both denominators of the integrand
(k2−m2+iε)((K −k)2−m2+iε)
vanish, namely when k2=m2and(K−k)2=m2. But we found the condition in exercise I.7.4, namely
thatK2≥4m2. Referring to (III.1.14)
M=iλ2
32π2/integraldisplay1
0dαlog/bracketleftbigg/Lambda12
α(1−α)K2−m2+iε/bracketrightbigg
we see that the log has a cut starting at K2=m2/α(1−α).A sα ranges from 0 to 1, the minimum value
ofm2/α(1−α)is attained at α=1
2. So indeed, the cut starts at K2=4m2.
552 | Solutions to Selected Exercises
III.1.3 Under the indicated change, log /Lambda1→logeε/Lambda1=log/Lambda1+ε, and so δM=−iδλ+iCλ23(2ε)+O(λ3).
Thus, δM=0 implies δλ=6Cλ2ε+O(λ3)=6Cλ2δlog/Lambda1+O(λ3)giving the stated result for
/Lambda1(dλ/d/Lambda1).
III.2.1 For/integraltext
ddx(∂ϕ)2to be dimensionless, we need [ϕ ]=(d−2)/2. Thus [ ϕn]=n(d−2)/2 and so in order for/integraltext
ddxλnϕnto be dimensionless, we must have [ λn]=n(2−d)/2+d.
III.3.3 When we set m=0, the integrand is manifestly a linear combination of γmatrices. The integral cannot
produce a term independent of the γmatrices, which is what Bis. For electrodynamics, the integral is
changed to
(ie)2i2/integraldisplayd4k
(2π)41
k2[(1−ξ)kμkν
k2−gμν]γμ/negationslashp+ /negationslashk+m
(p+k)2−m2γν
≡A(p2)/negationslashp+B(p2)
When m=0, the integrand is a linear combination of the product of three γmatrices, which can only
reduce to one γmatrix, not to none. Incidentally, an alternative way of seeing the results stated here is
to recall from chapter II.1 that with m=0 the Lagrangian is invariant under the chiral transformation
ψ→eiθγ5ψ.
III.3.4 This essentially follows from D=4−BE−3
2FEforBE=0 and FE=2. Then D=1 but the linear
divergence is reduced to logarithmic divergence by the symmetry argument given in the text.
III.5.2 Basically, you have already done this problem in exercises II.1.3 and II.1.4. You merely have to replace
Eand/vectorpby∂/∂t and/vector∇(see also chapter III.6).
III.5.3 In nonrelativistic quantum mechanics, the scattering amplitude is given in the Born approximation by
itimes the Fourier transform of the potential: i/integraltext
d3xei/vectork./vectorxU(/vectorx). The scattering amplitude owing to the
exchange of a scalar meson of mass mis just i/(k2−m2)/similarequal−i/(/vectork2+m2). Thus, we just repeat the
calculation in chapter I.4, obtaining
U(/vectorx)=−/integraldisplayd3k
(2π)3ei/vectork./vectorx
/vectork2+m2=−1
4πre−mr
III.6.1 ¯u(p/prime)(/negationslashp/primeγμ+γμ/negationslashp)u(p) =2m¯u(p/prime)γμu(p) by the equation of motion, but using γμγν=1
2{γμ,γν}+
1
2[γμ,γν]=ημν−iσμνwe can also write (/negationslashp/primeγμ+γμ/negationslashp)=(p/prime+p)μ+iσμν(p/prime−p)ν. We thus obtain
the Gordon decomposition.
III.6.2 We compute
qμ¯u(p/prime)[γμF1(q2)+iσμνqν
2mF2(q2)]u(p)=¯u(p/prime)/negationslashqu(p)F 1(q2)
=¯u(p/prime)(/negationslashp/prime− /negationslashp)u(p)F 1(q2)
=¯u(p/prime)(m−m)u(p)F 1(q2)=0
where the first and third equality follows from the antisymmetry of σand the equation of motion,
respectively.
III.7.1 Proceeding as in the text but living in d−dimensional spacetime, we obtain i/Pi1μν(q)=−i/integraltextddl
(2π)dNμν
D
where1
D=/integraltext1
0dα1
Dwith D=(l2+α(1−α)q2−m2+iε)2as before but with Nμνnow effectively equal
to−d(( 1−2
d)gμνl2+α(1−α)(2qμqν−gμνq2)−m2gμν). Rotating to Euclidean space we see that we
have to do the integrals (with c2≡m2−α(1−α)q2)/integraltextdd
El
(2π)d1
(l2+c2)2and/integraltextdd
El
(2π)dl2
(l2+c2)2=/integraltextdd
El
(2π)d1
(l2+c2)−
c2/integraltextdd
El
(2π)d1
(l2+c2)2. I did the first of these integrals for you in appendix II in chapter III.1. Generalizing
slightly we have/integraltext∞
0dlld−1 1
(l2+c2)a=1
2cd−2a/integraltext1
0dx( 1−x)d
2−1xa−1−d
2. I will let you carry on from here.
Solutions to Selected Exercises | 553
Part IV
IV .1.1 Write /vectorϕ=(ϕ1,ϕ2,..., v+ϕ/prime
N). We compute1
2μ2/vectorϕ2−(λ/4)(/vectorϕ2)2up toO(ϕ3)and find (upon dropping
the/prime)
1
2μ2(v2+2vϕN+/vectorϕ2)−λ
4(v4+4v2ϕ2
N+4v3ϕN+2v2/vectorϕ2)
The condition of no linear term in ϕNfixesv2=μ2/λand so the coefficient of /vectorϕ2is equal to1
2μ2−
λ/4(2v2)=0. The (N −1)fields ϕ1,ϕ2,..., ϕN−1are massless.
IV .3.1 We have/integraltext∞
0dklog[(k2+a2)/k2]=πaand so Veff(ϕ)=V( ϕ)+/planckover2pi√V/prime/prime(ϕ)/2 +O(/planckover2pi2).F o r L=1
2(∂ϕ)2−
1
2ω2ϕ2, the quantum oscillator with ϕidentified as position, we have Veff(0)=1
2/planckover2piω.
IV .3.3 We have
m(ϕ)=fϕ inVF(ϕ)=2i/integraltextd2p
(2π)2logp2−m(ϕ)2
p2
which after Wick rotation becomes
−2/integraldisplayd2pE
(2π)2logp2
E+m(ϕ)2
p2
E=−1
2π/integraldisplay∞
0dxlogx+m(ϕ)2
x
After cutting off the integral at /Lambda12and adding a counterterm Bϕ2we obtain
VF=1
2π(f ϕ)2logϕ2
M2
IV .3.4 Veff=i/summationtext∞
n=1/integraltext
d4k/(2π)4(1/2n)[V/prime/prime(ϕ)/k2]n.F o rV/prime/prime(ϕ)=1
2λϕ2the corresponding Feynman diagrams
consist of a circle with nV ’s attached to the circumference, where the 2 nis the infamous symmetry
factor that I tried to avoid talking about in chapter I.7.
IV .4.1 WithHan arbitrary p-form,
ddH=1
(p+1)!1
p!∂λ∂νHμ1μ2...μpdxλdxνdxμ1dxμ2...dxμp
=1
21
(p+1)!1
p![∂λ,∂ν]Hμ1μ2...μpdxλdxνdxμ1dxμ2...dxμp=0
IV .5.1 If you have done all the exercises thus far (see exercise I.10.3), you have already made the acquaintance of
theI=2 scalar field transforming as ϕab→RacRbdϕcd=Racϕcd(RT)db=(RϕRT)ab, which thus can
be written as a traceless 3 by 3 symmetric matrix ϕ→RϕRT. Now you merely have to write out the
covariant derivative Dμϕ(see IV .5.20) explicitly. [Hint: The action of the generators on ϕis similar to
what is shown in (B.20).]
IV .5.2 dF=d(dA+A2)=dAA−AdA and [A,F]=AdA−dAA and so dF+[A,F]=0. Explicitly with
indices, this reads εμνλσ(∂νFλσ+[Aν,Fλσ])=0. In the abelian case, we have, for μ=0,εijk∂iFjk=
/vector∇./vectorB=0 (recall chapter IV .4!), and for μ=i,εijk(−∂ 0Fjk+∂jF0k−∂jFk0)=−∂0Bi+(/vector∇×/vectorE)i=0.
IV .5.4 From the general arguments mentioned in the problem we know that tr F2must be the “d of some-
thing.” Now dtrAdA=trdAdA anddtr2
3A3=2
3tr(dAA2−AdAA +A2dA)=2t rdAA2but on the
other hand tr F2=tr(dA+A2)(dA+A2)=tr(dAdA +2dAA2)since tr A4=trA3A=−tr AA3=
−trA4=0. In electromagnetism, tr A3=0, and dtrAdA when written out in elementary notation
is just ∂μ(εμνλσAν∂λAσ)=1
4εμνλσFμνFλσ.
554 | Solutions to Selected Exercises
IV .5.6 We simply plug in the general expression in the text and obtain
L=−1
4g2Fa
μνFaμν+¯q(iγμDμ−m)q
with the covariant derivative Dμ=∂μ−iAμ=∂μ−iAa
μTa, where Ta(a=1 ,...,8 )are traceless her-
mitean 3 by 3 matrices. Explicitly, (A μq)α=Aa
μ(Ta)αβqβ, with α,β=1, 2, 3 (see chapter VII.3).
IV .6.3 Observe
Aa
μτa=/parenleftBiggA3
μA1−i2
μ
A1+i2
μ−A3μ/parenrightBigg
with the obvious notation A1±i2
μ≡A1
μ±iA2
μ. Let/angbracketleftϕ/angbracketright=/parenleftBig
0
v/parenrightBig
so that
Dμϕ=∂μϕ−i(gAa
μτa
2+g/primeBμ1
2)ϕ→−i
2v/parenleftBigggA1−i2
μ
−gA3
μ+g/primeBμ/parenrightBigg
Thus, |Dμϕ|2contains v2[g2A1+i2
μA1−i2
μ+(−gA3
μ+g/primeBμ)2]. The combinations A1+i2
μ,A1−i2
μ, and
(−gA3
μ+g/primeBμ)acquire mass while (g/primeA3
μ+gBμ)remain massless.
IV .7.4 We have
/Delta1μν(k1,k2)=i/integraldisplayd4p
(2π)4Nμν
D+{μ,k1↔ν,k2}
where
Nμν≡trγ5(/negationslashp− /negationslashq+M)γν(/negationslashp− /negationslashk1+M)γμ(/negationslashp+M)
Only the term linear in MinNμνdoes not vanish, giving Nμν=4iMεμνστk1σk2τ. Since we are interested
only in terms of O(k1k2)we can set D→(p2−M2)3so that
/Delta1μν(k1,k2)=− 8Mεμνστk1σk2τ/integraldisplayd4p
(2π)41
(p2−M2)3=i
4π2Mεμνστk1σk2τ
with a dependence on Mas stated in the problem. The effect of the regulator, like some unsavory
acquaintance, remains even after we have sent him to infinity.
IV .7.5 We will sketch the solution. The details may be found in the lectures given by S. Adler at the 1970
Brandeis Summer School. The point is to imagine a regularization scheme that preserves the variousrelevant symmetries, namely Lorentz invariance, vector current conservation, and Bose statistics. As you
will see, we don’t actually have to specify the regularization. By Lorentz invariance, we have
/Delta1λμν(k1,k2)=ελμνσk1σA1+ελμνσk2σA2+ελμστk1σk2τkν
1A3
+ελμστk1σk2τkν
2A4+ελνστk1σk2τkμ
1A5
+ελνστk1σk2τkμ
2A6+εμνστk1σk2τkλ
1A7
+εμνστk1σk2τkλ
2A8
Since the Feynman integral representing /Delta1λμνis superficially linearly divergent, we see that A3,...,A8
are all convergent since we have to pull out three powers of momentum to extract them. In contrast, A1
andA2are logarithmically divergent. But we can relate them to A3,...,A8by vector current conservation
since 0 =k1μ/Delta1λμν=ελνστk1σk2τ(−A2+k2
1A5+k1.k2A6)and thus A2=k2
1A5+k1.k2A6. Similarly for
A1. Rationalizing the Feynman integrand and evaluating the trace in the numerator, we can systematically
ignore terms that contribute only to A1andA2. Furthermore, Bose statistics gives us relations such as
A3(k2
1,k2
2,q2)=−A6(k2
2,k2
1,q2).
Solutions to Selected Exercises | 555
Part V
V .1.1 We dropped the term h2∂0θbut kept the term 4 g2¯ρh2. This requires ∂0θ/lessmuchg2¯ρ, that is, ω/lessmuchg2¯ρ, but
since in our solution ω∼g√¯ρ/mk this requires k/lessmuchg√m¯ρ, which is consistent with what we assumed
about k. Looking at the terms −2√¯ρh∂ 0θ−4g2¯ρh2inLwe see that h∼∂0θ/(g2√¯ρ)/lessmuch√¯ρ, which is
also consistent.
V .5.1 Withγ5=σ3,1
2(I±γ5)clearly projects out the top and bottom component of ψ=/parenleftBig
ψLψR/parenrightBig
, respectively.
Everything is formally the same as in chapter II.1, but we can also work things out explicitly in the
specific representation given here. Thus, ¯ψψ=ψ†σ2ψ=i(ψ†
RψL−ψ†
LψR)and¯ψγ5ψ=ψ†σ2σ3ψ=
i(ψ†
RψL+ψ†
LψR). Under the transformation ψ→eiθγ5ψ,ψL→eiθψLandψR→e−iθψR, and the
massless Dirac Lagrangian
L=iψ†
R(∂
∂t+vF∂
∂x)ψR+iψ†
L(∂
∂t−vF∂
∂x)ψL
clearly does not change.
V .6.1 This of course just follows from Lorentz invariance. We have
∂tϕ(x−vt√
1−v2)=−v√
1−v2ϕ/prime(x−vt√
1−v2)
and
∂xϕ(x−vt√
1−v2)=1√
1−v2ϕ/prime(x−vt√
1−v2)
and thus the equation
(∂2
t−∂2
x)ϕ(x−vt√
1−v2)+V/prime[ϕ(x−vt√
1−v2)]=0
becomes
ϕ/prime/prime(x−vt√
1−v2)−V/prime[ϕ(x−vt√
1−v2)]=0
Note that this does not depend on the form of V. For any relativistic theory, the soliton moves like a
relativistic particle (obviously!).
V .6.2 The sine-Gordon theory has an infinite number of vacua occurring at ϕ=(2n+1)π/β . Thus, there
exists a whole spectrum of solitons, such that ϕ(±∞)=(2n±+1)π/β . The topological current is
Jμ=(β/2π)εμν∂νϕwith the corresponding charge Q=(n+−n−). TheQ=2 soliton decays into two
Q=1 solitons.
V .7.4 (i/2π)/integraltext
S1gdg†=(i/2π)/integraltext
S1eiνθde−iνθ=(i/2π)/integraltext
S1(−iνdθ) =(ν/2π)/integraltext2π
0dθ=ν, which indeed counts
the number of times eiνθwinds around the circle. What mathematicians call the winding number is
indeed just the magnetic flux of the physicist.
V .7.5 Within a region small enough so that we can treat ϕa=vδa3as constant, using (Dμϕ)b=∂μϕb+
eεbcdAc
μϕdwe have (Dμϕ)1=evA2
μand(Dμϕ)2=−evA1
μand thus
Fμν≡Fa
μνϕa
|ϕ|−(1/e)εabcϕa(Dμϕ)b(Dνϕ)c
|ϕ|3
→F3
μν+e(A2
μA1ν−A2
νA1μ)=∂μA3
ν−∂νA3
μ
precisely the electromagnetic field strength since A3
μis the massless component of the Yang-Mills field.
Let us compute Bk=εijkFijfar from the magnetic monopole. To calculate the magnetic charge we are
556 | Solutions to Selected Exercises
interested only in the term of order 1 /r2in/vectorB. Since Dμϕ→O(1/r2)by construction we can drop the
second term in Fij. Thus, we merely have to compute Fa
ij≡∂iAa
j−∂jAa
i+eεabcAb
iAcj. Since Fa
ijwill
eventually be contracted with the unit vector ϕa/|ϕ|=xa/r, we can effectively drop some of the terms
inFa
ij, thus simplifying the computation. We have
∂iAa
j=∂i(1
eεajlxl
r2)“="1
eεaji1
r2
and
eεabcAb
iAcj=(1/e)εabcεbimεcj nxmxn
r4
=(1/e)(δciδam−δcmδai)εcj nxmxn
r4=1
er4εijnxaxn
so that
Fa
ijϕa
|ϕ|=Fa
ijxa
r=1
er3(−2+1)εaijxa=−1
er3εaijxa
and hence Bk=−(1/er2)ˆxk. The magnetic charge g=− 4π/e .
Our result appears to differ from Dirac’s quantization condition (IV .4.10) by a factor of 2. The
resolution of this apparent paradox is instructive. In fact, we can always introduce into this theory a
field/Psi1(which could be a Bose or a Fermi field) transforming in the I=1
2representation with the
corresponding covariant derivative Dμ/Psi1=∂μ/Psi1−ie(1
2τa)Aa
μ/Psi1. The field /Psi1carries electric charge1
2e.
Thus, the fundamental unit of electric charge is actually1
2e, note, and our result g=− 4π/e=− 2π/(e/2 )
is actually nothing but the Dirac quanization condition. (The sign is trivial: just a question of which onewe call the monopole and which the antimonopole.)
V .7.7 Plugging in the Ansatz ϕ
a=(H(r)/er)(xa/r) andAb
i=[1−K(r) ]εbij(xj/er2)[so that H(r)−→
r→∞evr and
K(r)−→
r→∞0 in accordance with the asymptotic behavior (V .7.5) and (V .7.6)] into M=/integraltext
d3x{1
4(/vectorFij)2+
1
2(Di/vectorϕ)2+V(/vectorϕ)}we get Mas a functional of HandK. Minimizing Mgives (with H/prime=dH/dr etc) the
equations r2H/prime/prime=2HK2+(λ/e2)[H3−(ev)2r2H] andr2K/prime/prime=K(K2−1)+KH2. For help, see M. K.
Prasad and C. M. Sommerfeld, Phys. Rev. Lett. 35: 760, 1975.
V .7.8 The BPS solution corresponds to setting λ=0 in the two equations in exercise V .7.7, rendering the
equations soluble, with the solution H(r)=evr( cothevr)−1 andK(r)=evr/( sinhevr) . Ask yourself
whyH(r) andK(r) approach their asymptotic values exponentially. What determines the length scale?
V .7.9 For help, see B. Julia and A. Zee, Phys. Rev. D11: 2227, 1975.
V .7.11 We derived the lower bound for the mass of the magnetic monopole 4 πv|g|∼4π(ev)/e2∼MW/α.
V .7.12 Near the identity element g=ei/vectorθ./vectorσ/similarequal1+i/vectorθ./vectorσand thus gdg†/similarequal−id/vectorθ./vectorσ. In a small neighborhood of
the identity element the group manifold is locally Euclidean and so
tr(gdg†)3=itr(σiσjσk)dθidθjdθk=− 12dθ1dθ2dθ3
is manifestly proportional to the volume element on S3.F o r g=ei(θ1σ1+θ 2σ2+mθ 3σ3),t r(gdg†)3=
−12mdθ1dθ2dθ3.
V .7.13/integraltext
d4x(∂μJμ
5)=/integraltext
d3xJ0
5|t=+∞−/integraltext
d3xJ0
5|t=−∞ . Recalling that J0
5=ψ†
RψR−ψ†
LψL, we see that the two
spatial integrals just count the number of right moving fermion quanta minus the number of left moving
fermion quanta at t=± ∞ respectively. So/integraltext
d4x(∂μJμ
5)is an integer. On the other hand, in the text we
proved that/integraltext
trF2is a topological invariant. In other words, with suitable normalization, evidently
1/(4π)2,the integral [1 /(4π)2]/integraltext
d4xεμνλσtrFμνFλσis an integer. Thus, the coefficient 1 /(4π)2cannot
be shifted even a little bit by quantum fluctuations.
Solutions to Selected Exercises | 557
Part VI
VI.4.2 The quartic interaction term (1/2f2)(/vectorπ.∂/vectorπ)2inLgives the amplitude i(1/2f2)i2δabδcd(k1k3+k1k4+
k2k3+k2k4)+permutations =(i/2f2)δabδcd(k1+k2)2+permutations for the 4-pion interaction vertex
(where for convenience we have labeled all the momenta as going outward so that k1+k2+k3+k4=0).
VI.4.3 After writing σ=v+σ/prime, we find, as in chapter IV .1, that L=−1
2(2μ2)σ/prime2−λvσ/prime/vectorπ2−1
4λ(/vectorπ2)2+
... , where we have displayed only terms relevant for our purposes. The diagrams contributing to
four-pion interaction are of two types, those involving the λ(/vectorπ2)2term and those invoking σ/primeex-
change. The former gives for the amplitude (−i
4λ)2 .2(δabδcd+δacδbd+δadδbc)while the latter gives
(−iλv)2{2i/[(k1+k2)2−m2
σ/prime]}δabδcd. Thus, expanding to first order in momenta squared we find the
coefficient of δabδcd:
−iλ−2iλ2v2/bracketleftBigg
−1
m2
σ/prime/bracketrightBigg
(1+(k1+k2)2
m2
σ/prime)=−iλ+2iλ2v2
2μ2/bracketleftbigg
1+(k1+k2)2
2μ2/bracketrightbigg
=iλ
2μ2(k1+k2)2
To compare with exercise VI.4.2 we remember that f2=v2=μ2/λso that the amplitude here is also
equal to (i/2f2)δabδcd(k1+k2)2+permutations, as we had anticipated in the text.
VI.4.4 We will track down factors of 2 carefully but not factors of iand−1. Let us go back to the chiral
transformations ψ→[1+i/vectorθ.(/vectorτ/2)γ5]ψand¯ψ→¯ψ[1+i/vectorθ.(/vectorτ/2)γ5]. Thus, δ(¯ψψ)=θa¯ψiγ5τaψand
δ(¯ψiγ5τaψ)=−θa¯ψψ . Hence, for L=¯ψ{iγ∂+g(σ+i/vectorτ./vectorπγ 5)}ψ+L(σ,/vectorπ)to be invariant we must
haveδσ=θaπaandδπa=−θaσ. Applying Noether’s theorem Jμ=(δL/δ∂μϕ)δϕ, we obtain the current
Ja
μ5=¯ψiγμγ5(τa/2)ψ+πa∂μσ−σ∂μπawritten in the text. Comparing the term ¯piγμγ5ncontained in
J1+i2
μ5≡J1
μ5+iJ2
μ5with the current J5μdefined in chapter IV .2, we see that J5μ=−iJ1+i2
μ5. The normal-
ized state |π−/angbracketright=(1/√
2)(|π1/angbracketright−i|π2/angbracketright)so that /angbracketleft0|π1+i2|π−/angbracketright= 2/√
2. The current J1+i2
μ5contains the
term−v∂μπ1+i2and thus f=√
2v. Next, we have to work out the pion-nucleon coupling gπNN as de-
fined in chapter IV .2. Here Lcontains g¯ψi/vectorτ./vectorπγ5ψ, which contains√
2g¯piγ 5nπ−sinceπ1−i2=√
2π−.
Thus, gπNN=√
2g. Putting it together, we see that M=gvtranslates to 2 M=fgπNN in agreement
with chapter IV .2.
VI.6.1 See figure VI.6.1. From /Delta1h=(d/cos θ)/similarequald(1+1
2θ2)and(∂h/∂x) =tanθ/similarequalθ, we have (∂h/∂t) ∝θ2∝
(∂h/∂x)2, thus giving rise to the term (λ/2)(/vector∇h)2. It all goes back to Mr. Pythagoras.
VI.6.2 We integrate the term1
2/integraltext
dD/vectorxd t [((∂/∂t) −/vector∇2)h]2inS(h) by parts to obtain −1
2/integraltext
dD/vectorxd t [h((∂/∂t) +
/vector∇2)((∂/∂t) −/vector∇2)h]. Thus, the propagator is the inverse of the operator (∂/∂t +/vector∇2)(∂/∂t −/vector∇2)=
∂2/∂t2−(/vector∇2)2, the Fourier transform of which is −(ω2+k4).
VI.8.3 Withh/prime(/vectorx,t)=h/parenleftbig
/vectorx+g/vectorut,t/parenrightbig
+/vectoru./vectorx+g
2u2t, we have ∂h/prime/∂t=∂h/∂t +g/vectoru./vector∇h+(g/2)u2and/vector∇h/prime=
/vector∇h+/vectoru. Thus the combination (∂h/∂t) −g
2(/vector∇h)2is invariant, as is (obviously) /vector∇2h. In other words,
˜S(h) must be constructed out of these two invariant combinations.
VI.8.5 Look at the action S(h)=1
2/integraltext
dD/vectorxd t [(∂h/∂t) −∇2h−g(∇h)2/2]2. Comparing ∂h/∂t −∇2hwe see that
time has the dimension of length squared: T∼L2. From the term/integraltext
dD/vectorx dt (∂h/∂t)2and the fact that
Sis dimensionless, we have [h]2∼T2/(LDT)∼1/LD−2and so hhas the dimension of (1/LD−2)1
2.
Comparing /vector∇2hwithg(/vector∇h)2we see that ghas the dimension of 1 /h, that is, L(D−2)/2.
VI.8.7 We are told that L(dg/dL) =(2−D)g/ 2+(2D−3)fDg3+... . We are assuming that the terms
(...)can be neglected. Thus (in what follows a2andb2are two generic positive numbers) for D=1,
L(dg/dL) =a2g−b2g3andgflows toward the fixed point g∗=a/b . (Incidentally, the KPZ equation
is soluble for D=1 by methods not explained in this text and both zandχare known exactly.) For
D=2,L(dg/dL) =b2g3andgflows toward some unknown strong (presumably) coupling fixed point.
ForD=3,L(dg/dL) =−a2g+b2g3. The fixed point g∗=a/b is unstable. For g<g∗,gflows toward
the trivial (i.e., free, or Gaussian) fixed point. Since the theory at the fixed point is free we know the
558 | Solutions to Selected Exercises
critical exponents exactly: z=2 and χ=(2−D)/ 2. For g>g∗,gflows toward some unknown strong
(presumably) coupling fixed point.
Part VII
VII.1.1. SetnμA/prime
μ(x)=0 with A/prime
μ=U†AμU+iU†∂μU, so that n.∂U(x) =in.A(x)U(x). Define λ(x)=
r.x/(r .n)for any 4-vector rand write x=λ(x)n +x⊥, so that r.x⊥=0. Then
U(x)=Pei/integraltextλ(x)
0dσn.A(σn+x ⊥)
(with a path ordering) solves the differential equation, since n.∂λ=1 by construction.
VII.1.2 Using the BHC formula given, we have (the V’s are clearly irrelevant)
UijUjk=eiaAμeiaAν=eia(Aμ+Aν)−1
2a2[Aμ,Aν]+a3C+a4D+O(a5)
Similarly,
UklUli=e−iaA/prime
μe−iaA/primeν=e−ia(A/prime
μ+A/primeν)−1
2a2[Aμ,Aν]+a3E+a4F+O(a5):
the prime reminding us that the AμandAνin this expression is to be evaluated on the “north” and
“west” side of the plaquette in figure VII.1.2, respectively, in contrast to the AμandAνinUijUjkwhich are
evaluated on the “south” and “east” side, respectively. Here C,D,E, andFdenote various commutators,
which we drag along merely to show that they eventually drop out in what interests us. (Note how thedifferent terms are associated with different powers of a, as indicated. Note also that in some places we
have dropped the prime on Aand absorbed the “error” in doing so into terms of higher order in a.) Thus,
UklUli=e−ia(A μ+Aν)−ia2(∂νAμ−∂μAν−1
2i[Aμ,Aν])+a3G+a4H+O(a5)
where GandHdenote sums of commutators and terms such as ∂ν∂νAμand∂ν∂ν∂νAμ. Applying the
BHC formula again to the order indicated we have
UijUjkUklUli=eia2(∂μAν−∂νAμ)−a2[Aμ,Aν]+O(a4)=eia2Fμν+a3I+a4J+O(a5)
withFμν=∂μAν−∂νAμ+i[Aμ,Aν]. The same remarks on GandHapply to IandJ. The Yang-Mills
field strength emerges naturally, as we would anticipate. Since the traces of commutators and of Avanish,
when we apply the trace all the junk drops out to O(a5)and we have
S(P)=Re tr[1 −1
2a4FμνFμν+O(a5)]
By gauge invariance, the corrections must be of even order in abut for our purposes we don’t care about
them anyway. Evidently, fandgare related by some uninteresting factors of a.
Part VIII
VIII.1.7 R12=dω12=d(−cosθdϕ)=sinθdθdϕ =2R12
θϕdθdϕ /equal1⇒R12
θϕ=1
2sinθ. Since eθ
1=1,eϕ
2=1/sinθ,
we obtain R≡Rab
μνeμ
aeν
b=2R12
θϕeθ
1eϕ
2=1, independent of θandϕas expected.
Part N
N.1.1. The effective action for an electrically neutral system is given in the point particle limit by S=/integraltext
dτ(−m+
bEEμEμ+bBBμBμ+...), with EμandBμdefined in the text. The interaction terms involve two powers
of derivatives, which translate into two powers of ωin the scattering amplitude and hence four powers
ofωin the scattering cross section. (Note that a possible term like/integraltext
dτFμνFμνcan be absorbed into the
two terms already displayed.)
N.3.2. As in (III.3.7) we have 3 V3+4V4=2I+nwhere Idenotes the number of internal lines. The number
of loops (III.3.6) L=I−(V3+V4−1)is 0 in a tree diagram. Thus V3=n−2−2V4≤n−2.
Further Reading
Books on field theory
This is a list of field theory textbooks that I know about. I do not necessarily recommend them all.
In food as in books, each has his or her own taste.
T . Banks, Modern Quantum Field Theory, Cambridge University Press, New York, 2008.
J. D. Bjorken and S. D. Drell, Relativistic Quantum Mechanics , McGraw-Hill, New York, 1964.
———, Relativistic Quantum Fields , McGraw-Hill, New York, 1965.
L. S. Brown, Quantum Field Theory , Cambridge University Press, New York, 1992.
S. J. Chang, Introduction to Quantum Field Theory , World Scientific, Singapore, 1990.
T . P. Cheng and L. F. Li, Gauge Theory of Elementary Particle Physics , Clarendon Press, Oxford, 1984.
F. Dyson and D. Derbes, Advanced Quantum Mechanics, World Scientific, Singapore, 2007.
R. P. Feynman, Quantum Electrodynamics, W . A. Benjamin, New York, 1962.
K. Huang, Quantum Field Theory , John Wiley & Sons, New York, 1998.
C. Itzykson and J-B. Zuber, Quantum Field Theory , McGraw-Hill, New York,1980.
T . D. Lee, Particle Physics and Introduction to Field Theory , Taylor & Francis, New York, 1981.
V . P. Nair, Quantum Field Theory, Springer, New York, 2005.
M. E. Peskin and D. V . Schroeder, An Introduction to Quantum Field Theory , Perseus, Reading MA,
1995.
L. H. Ryder, Quantum Field Theory , 2nd Ed., Cambridge University Press, New York, 1996.
M. Stednicki, Quantum Field Theory, Cambridge University Press, New York, 2007.
G. Sterman, An Introduction to Quantum Field Theory, Cambridge University Press, New York, 1993.
S. Weinberg, Quantum Theory of Fields ,V o l s .1&2 ,C ambridge University Press, New York, 1996.
X. G. Wen, Quantum Field Theory of Many-Body Systems, Oxford University Press, New York, 2007.
and finally, of course,
F. Mandl, Introduction to Quantum Field Theory , Interscience, New York, 1959.
560 | Further Reading
Books on various topics mentioned
A. A. Abrikosov, L. Gorkov, and A. Dzyaloshinski, Methods of Quantum Field Theory in Statistical
Physics , Prentice Hall, Englewood Cliffs, NJ, 1963.
S. L. Adler, “Perturbation Theory Anomalies,” in: Lectures on Elementary Particles and Quantum Field
Theory , 1970, Brandeis University Summer Institute in Theoretical Physics, S. Deser et al, ed.,
MIT Press, Cambridge, 1970.
P. Anderson, Basic Notions of Condensed Matter Physics , Benjamin-Cummings, Menlo Park, CA 1984.
D. Bailin and A. Love, Supersymmetric Gauge Field Theory and String Theory , IOP Publishing, Bristol
and Philadelphia, 1994.
R. Balian and J. Zinn-Justin, eds., Methods in Field Theory , North Holland Publishing, Amsterdam,
and World Scientific, Singapore, 1981.
A. L. Barabasi and H. E. Stanley, Fractal Concepts in Surface Growth , Cambridge University Press,
Cambridge, 1995.
D. Budker, S. J. Freedman, and P. H. Bucksbaum, eds., Art and Symmetry in Experimental Physics:
Festschrift for Eugene D. Commins , American Institute of Physics, New York, 2001.
J. Cardy, Scaling and Renormalization in Statistical Physics , Cambridge University Press, New York,
1996.
S. Coleman, Aspects of Symmetry , Cambridge University Press, Cambridge, 1985.
J. Collins, Renormalization , Cambridge University Press, Cambridge, 1985.
E. D. Commins, Weak Interactions , McGraw-Hill, New York, 1973.
E. D. Commins and P. H. Bucksbaum, Weak Interactions of Leptons and Quarks , Cambridge University
Press, Cambridge, 2000.
M. Creutz, Quarks, Gluons and Lattices , Cambridge University Press, Cambridge, 1983.
P. A. M. Dirac, The Principles of Quantum Mechanics , Oxford University Press, Oxford, 1935. (On
p. 253 he explained why he wanted the equation of motion for the electron to be first order in timederivative.)
A. Dobado et al., Effective Lagrangians for the Standard Model , Springer-Verlag, Berlin, 1997.
O.J.P. ´Eboli et al., Particle Physics , World Scientific, Singapore 1992.
R. P. Feynman, Statistical Mechanics , Perseus Publishing, Reading, MA, 1998.
R. P. Feynman and A. R. Hibbs, Quantum Mechanics and Path Integrals , McGraw-Hill, New York,
1965.
J. M. Figueroa-O’ Farrill, Electromagnetic Duality for Children , on the World Wide Web 1998.
V . Fitch et al., eds., Critical Problems in Physics , Princeton University Press, Princeton, 1997.
M. Gell-Mann and Y . Ne’eman, The Eightfold Way , W . A. Benjamin, New York, 1964.
H. B. Geyer, ed., Field Theory, T opology and Condensed Matter Physics , Springer, 1995 (A. Zee,
“Quantum Hall Fluids.”)
M. L. Goldberger and K. M. Watson, Collision Theory, Dover, New York, 2004.
N. Goldenfeld, Lectures on Phase T ransitions and the Renormalization Group , Addison-Wesley, Read-
ing, MA, 1992.
F. Guerra and N. Robotti, Ettore Majorana: Aspects of His Scientific and Academic Activity, Springer,
New York, 2008.
C. Itzykson and J-M. Drouffe, Statistical Field Theory , Cambridge University Press, Cambridge, 1989.
S. Iyanaga and Y . Kawada, eds., Encyclopedic Dictionary of Mathematics , MIT Press, Cambridge, 1980.
L. Kadanoff, Statistical Physics , World Scientific, Singapore, 2000.
G. Kane and M. Shifman, eds., The Supersymmetric World: The Beginning of the Theory ,World
Scientific, Singapore, 2000.
J. I. Kapusta, Finite-T emperature Field Theory , Cambridge University Press, Cambridge, 1989.
L. D. Landau and E. M. Lifschitz, Statistical Physics , Addison-Wesley, Reading, MA, 1974.
S. K. Ma, Modern Theory of Critical Phenomena , Benjamin/Cummings, Reading, MA, 1976.
Further Reading | 561
H. J. W . M ¨uller-Kirsten and A. Wiedemann, Supersymmetry , World Scientific, Singapore 1987.
T . Muta, Foundations of Quantum Chromodynamics , World Scientific, Singapore, 1998.
D. I. Olive and P. C. West, eds., Duality and Supersymmetric Theories , Cambridge University Press,
Cambridge, 1999.
J. Polchinski, String Theory , Cambridge University Press, Cambridge, 1998.
J. J. Sakurai, Invariance Principles and Elementary Particles , Princeton University Press, Princeton,
1964.
L. Schulman, T echniques and Applications of Path Integrals , John Wiley & Sons, New York, 1981.
R. F. Streater and A. S. Wightman, PCT , Spin Statistics, and All That , W . B. Benjamin, New York,
1968.
G. ’t Hooft, Under the Spell of the Gauge Principle , Word Scientific, Singapore, 1994.
G. ’t Hooft et al., eds. Recent Developments in Gauge Theories , Plenum, New York, 1980.
D. Voiculescu, ed., Free Probability Theory , American Mathematical Society, Providence, R.I., 1997.
S. Weinberg, Gravitation and Cosmology , John Wiley & Sons, New York, 1972.
C. N. Yang, Selected Papers 1945–1980 with Commentary , W . H. Freeman, San Francisco, 1983.
A. Zee, Unity of Forces in the Universe , World Scientific, Singapore, 1982.
J.-B. Zuber, ed., Mathematical Beauty of Physics , World Scientific, Singapore, 1997.
Some popular books and books on the history of quantum field theory
M. Bartusiak, Einstein ’s Unfinished Symphony, Joseph Henry Press, Washington, D.C., 2000.
I. Duck and E. C. G. Sudarshan, Pauli and the Spin-Statistics Theorem , World Scientific, Singapore
1997.
R. P. Feynman, QED: The Strange Theory of Light and Matter, Princeton University Press, Princeton,
2006.
D. Kaiser, Drawing Theories Apart, University of Chicago Press, Chicago, 2005.
A. I. Miller, Early Quantum Electrodynamics , Cambridge University Press, Cambridge, 1994.
L. O’Raifeartaigh, The Dawning of Gauge Theory , Princeton University Press, Princeton, 1997.
S. S. Schweber, QED and the Men Who Made It: Dyson, Feynman, Schwinger, and T omonaga , Princeton
University Press, Princeton, 1994.
A. Zee, Fearful Symmetry , Princeton University Press, Princeton, 1999.
———, Einstein ’s Universe , Oxford University Press, New York, 2001.
———, Swallowing Clouds , University of Washington Press, Seattle, 2002.
Further Reading for Part N
In writing a textbook, the author has the luxury of not preparing a detailed scholarly bibliography
(unless he or she chooses to follow the example of S. Weinberg, who is, in my opinion, mostadmirable in this regard). Even more extravagant is the freedom accorded to authors of popularbooks who in most cases give their unsuspecting and gullible readers the impression that the physicsof an entire era was done by two or three greats, individuals worthy of their own personality cults.Presenting recent developments still in flux, I am faced with the dilemma of whether to give propercredit. In scholarly publications, conscientious referencing is of course ethically mandated, but thisis a textbook. Fortunately, in this age of omniscient search engines, the reader could easily compilea bibliography more exhaustive than even a myopic humanist used to be able to muster in half alifetime. I could do the same, but it is of little help to you for me to merely list the names of those
562 | Further Reading
responsible for, say, the new way of computing amplitudes using the spinor helicity formalism.1
Instead, I can best serve the typical reader by listing a few papers and review articles starting from
which you can track down the literature to your scholarly heart’s desire. To those who feel that theyshould be mentioned, I apologize and refer you to Glashow’s description of a tapestry in the preface.
W . Goldberger and I. Z. Rothstein, arXiv: hep-th/0409156v2.
Z. Bern, L. J. Dixon, D. C. Dunbar, D. A. Kosower, arXiv: hep-ph/9602280.N. Arkani-Hamed and J. Kaplan, arXiv: hep-th/0801.2385.E. Witten, arXiv: hep-th/0312171.
1F. A. Berends, Z. Bern, L. Chang, P. De Causmaecker, L. J. Dixon, D. C. Dunbar, R. Gastmans, W . Giele,
J. F. Gunion, R. Kleiss, D. A. Kosower, Z. Kunszt, M. Mangano, A. G. Morgan, S. J. Parke, W . J. Stirling, T . R.Taylor, W . Troost, T . T . Wu, Z. Xu, D. H. Zhang, and many many others. I know how to copy and paste also! Please
forgive me if I inadvertently left you off this list.
Index
Page numbers followed by letters f and n refer to figures and notes, respectively.
Abrahams, Elihu, 366
accelerators, 42, 483Adler, Steve, 277Aharonov-Bohm effect, 251–252, 261, 317, 320amplitudes: as analytic functions, 208–209;
symmetry in, 78. See also meson-meson scattering
amplitude
“amputating the external legs,” 55analyticity, in quantum field theory, 207–209, 211,
217, 219. See also nonanalyticity
Anderson, Phil: in “Gang of Four,” 366; Nobel Prize
for, 351
Anderson localization, 351, 354; in renormalization
group language, 366–367
Anderson mechanism, 264angular momentum, addition of, 530anharmonicity, in field theory, 43, 89anomaly (axial anomaly/chiral anomaly), 270, 275;
alternative ways of deriving, 279; consequences of,275–278; Feynman diagram calculation revealing,270–274; grand unification and freedom from,411–412, 429; higher-order quantum fluctuationsand, 310; in nonabelian gauge theory, 276;nonrenormalization of, 277; and path integralformalism, 278
anthropic selection, 451anticommutation: spin-statistics connection and,
122, 123; wave function of electrons and, 107
antiferromagnet(s): effective low energy descriptionof, 344–345; magnetic moments in, 344; N ´eel
state for, 346
antikink(s), 304, 305fantimatter: discovery of, 101; requirement of, 157antineutrino field, in SO (10) unification, 425–426
antiunitary operator, time reversal as, 103anyon(s), 315; interchanging, 316; statistics between,
317
approximation, steepest-descent, 16area law, 377, 387asymptotically free theories, 360, 386, 390;
Gross-Neveu model as, 403
atom(s), interaction with radiation, 3attraction: quantum field theory on, 35–36; spin 1
particle and, 36–37; spin 2 particle and, 36
auxiliary field, 192, 467–468axial anomaly. See anomaly
axial current conservation, quantum fluctuations
destroying, 274–275
axial gauge, 378
background field method, 504–507
Bardeen, Bill, 277bare perturbation theory, 175baryon number conservation, law of, 413; grand
unification and violation of, 418
BCFW recursion, 500, 507, 514
Berends, F. A., 493Bern, Zvi, 484, 484f, 516, 519
564 | Index
Bern transformation, 516
Berry’s phase, nonabelian, 261, 346Berry’s phase term, in ferromagnets and
antiferromagnets, 345, 346
beta decay, 456Bethe, Hans, 365Bianchi identity, 247, 261Bjorken, James, 356black hole(s): gravitational waves in, 479–482;
Hawking radiation from, 290–291; Schwarzschild,311
Bludman, Sid, Yang-Mills theory and, 379blue sky, effective field theory of, 457–458Bogoliubov, N. N., 192Bogoliubov calculation, of gapless mode, 284Bogomol’nyi inequality, 305; for mass of monopole,
309
Bogomol’nyi-Prasad-Sommerfeld (BPS) states, 309Bohr, Niels, 252Boltzmann, Ludwig, 150, 287; on entropy, 311Bose-Einstein condensation, 295Bose-Einstein statistics, 120Bose field, mass correction to, divergence of, 180boson(s): “bad” behavior of, 180; electron pairing
into, 295; and fermions, unification of, 461;gapless mode in, 284; gauge (see gauge boson[s]);
intermediate vector, and Fermi theory of theweak interaction, 171–172, 309; Lorentz invariantscalar field theory on, 190; mass correction for,divergence of, 180; massless, emergence of, 226–227; Nambu-Goldstone ( seeNambu-Goldstone
boson[s]); nonabelian gauge (Yang-Mills), 257,386, 434; in nonrelativistic theory, 192–193;repulsion of, 283, 337
BPS (Bogomol’nyi-Prasad-Sommerfeld) states, 309brane world scenarios, 40–42, 450Br´ezin, Edouard, 402
Brillouin zone, 298Brink, Lars, 470Britto, Ruth, 500Burgoyne, N., 121
Cachazo, F. A., 500
canonical formalism, 61–69; and degrees of freedom,
67; and Feynman diagrams, 43; vs. path integralformalism, 44, 61, 67; propagator in, 67–68; timeordering in, 67–68
Carrasco, J. J., 519Casimir force, between two plates, 70–75, 71fCauchy’s theorem, 209central identity of quantum field theory, 182, 524
chain rule, 445charge: in dual theory, vortices as, 332–334; asgenerator, 80; of quasiparticles, 327. See also
electric charge; magnetic charge
charge conjugation, 101–102; in grand unification,
429
charge quantization, deducing, 121Chern-Simons term, 317; gauge invariance of, 328;
for Hall fluid, 326; massive Dirac fermions and,319–320; and Maxwell term, 320; in nonrelativistictheory, 320
Chern-Simons theory, 318–319; effective theory of
Hall fluid as, 324
Chew, Geoff, 105chiral anomaly. See anomaly
chiral superfield, 464, 466chiral symmetry: condition for, 419; conserved
current associated with, 100; of strong interaction,
234, 387–388
classical limit, path integral formalism for taking, 19classical physics, symmetry of, 270Clifford algebra: and Dirac bilinears, 97; and Dirac
equation, 94–95
closed forms, 247coherence length, 296Coleman, Sidney, 252, 473Coleman-Mermin-Wagner theorem, 230Coleman-Weinberg effective potential, 240color, quark, 385, 386complex plane, 207, 208Compton, Arthur, 154Compton scattering, 152–157condensation, and superconductivity, 295condensed matter physics: critical dimension in,
364, 367; disordered systems studied by, 350, 354;goal of, 328; Goldstone’s theorem in, 229–230;impurities studied by, 354; length scales in, 169;momentum density in, 191; number conjugateto phase angle in, 192; particle physics and, 281,452–453; quantum field theory and, 5, 190, 281;quantum Hall effect and, 351–352; quasiparticlesin, 326; renormalization group in, 360–363;spin-statistics rule and, 120
conductivity, vs. conductance, 366–367connected graphs, vs. disconnected graphs, 29, 47conserved current: charge associated with, 80; and
chiral symmetry, 100; and continuous symmetry,78–79; momentum space version of, 133
continuous symmetry, 77–78, 226; conserved current
and, 78–79
continuous symmetry breaking, 226; Coleman-
Mermin-Wagner theorem on, 230; and masslessfields, 228–229
Cooper pairs of electrons, 295coordinate transformations, 81–82
Index | 565
cosmic coincidence problem, 450
cosmological constant, 448–449; measured in units
of Gev4, 449–450; order of magnitude, 449
cosmological constant problem, 450; approach
to, 455; root of, 448; string theory’s inability toresolve, 450; supersymmetry as solution to, 461
Coulomb potential, 133–134, 143; modification of,
205
Coulomb’s electric force: and Newton’s gravitational
force, comparison of, 29; quantum field theory on,32–33
counterterms: cutoff dependence absorbed by,
241; in Feynman diagrams, 175–176, 176f;nonrenormalizable theories and, 179, 241
coupling, electromagnetic, 358coupling constant(s), 164, 173; dimensionless,
170; of electromagnetic interaction, 359; hadronproliferation and, 231; as misnomer, 358–359;pion-nucleon, 235; renormalized, 166; Yang-Mills,258–259
coupling renormalization, 173–174CPT theorem, 104critical dimension, in condensed matter physics,
364, 367
critical phenomena, 292; complete theory of, 293;
Landau-Ginzburg theory of, 293–294
crossing, 156cubic vertex, in spinor helicity formalism, 496current conservation. See conserved current
curved spacetime: Dirac action in, 445; introduction
to, 84–86; quantum field theory in, 82, 290
Cutkosky cutting rule, 215–216, 217, 219, 493cutoff dependence, 163; avoiding in physical
perturbation theory, 176; counterterms absorbing,241; disappearance of, 166, 167; of meson-mesonscattering amplitude, 173
dark energy, 450
dark matter, 450Dashen, Roger, 405decay rate, 139–141, 212Deser, Stanley, 470differential forms, 246–247; use in nonabelian gauge
theory, 255, 256
differential operator, propagator as inverse of, 23dilatation invariance, 84ndimensional analysis, 169–170, 453; on meson-
meson scattering amplitude, 173
dimensional regularization, 167, 168, 204Dirac, Paul, 105; on electric charge, quantizing of,
410; on magnetic monopole, 308; metaphorical
language of, 113; on path integral formalism,10–13; on positron, 5; on quantum mechanicsand magnetic monopoles, 245; on spinor
representation, 117; teaching style of, 454
Dirac basis vs. Weyl basis, 98–99Dirac bilinears, 97Dirac equation, 93–105; Clifford algebra and, 94–
95; in curved spacetime, 444; and degrees offreedom, reduction in, 95; derivation of, 118;electromagnetic field and, 101; handedness and,100; Lorentz transformation and, 96–97; magneticmoment of electron in, 194–195; origins of, 93–94; parity and, 98; in solid state physics, 298, 299;solving, 98; time reversal and, 104
Dirac field: interacting with scalar field, Feynman
rules for, 53–54, 534–535; interacting with vectorfield, 100; interacting with vector field, Feynmanrules for, 129, 129f, 535–536; propagator for,
127; quantizing, 107–113, 122; quantizing byGrassmann path integral, 127; vacuum energy of,111–112, 125
Dirac operator, 113Dirac spinor, 94, 96; components of, 117; and
supersymmetry, 114
disorder: Anderson localization of, 351, 354;
condensed matter physics and study of, 350, 354;Grassmannian approach to, 354
dispersion relations, 208–210, 217–218, 235Di Vecchia, Paolo, 470divergence(s): degree of, 176–178; dependence on
dimension of spacetime, 179; with fermions, 178–179; logarithmic, 175, 176–177; in quantum fieldtheory, 57–58, 161–162; superficial degree of,176, 179; total, supersymmetric transformation,465–466
dotted and undotted notation, 116–117, 541–544;
replacing, 475
double-line formalism, 395, 396double-slit experiment, 7, 8f; expansion of, 7–9, 8f, 9fdouble-well potential, 224, 224fDrell, Sid, 356duality: in (2+1)-dimensional spacetime, 335;
concept of, 331, 332; electromagnetic, 249; andlinking of perturbative weak coupling to strongcoupling, 473; of monopoles, 334; nonrelativistictreatment of, 336–337; relativistic treatment of,335–336; of string theories, 334; vortex, 334
dual theory, vortices as charges in, 332–334dynamical symmetry breaking, 230; example of, 388dynamical variable, in quantum field theory, 19dyon, 309Dyson, Freeman, 60Dyson gas approach, 400–402
effective field theory: of blue sky, 457–458;
566 | Index
effective field theory (continued)
development of, 452; Fermi theory of the weakinteraction as, 456; gravitational waves and, 479–482; of Hall fluid, 324–325, 452–453; of neutrinomasses, 456; predictive power of, 456–459; ofproton decay, 455–457; recent developments in,479–482; and renormalization group flow, 453;reshuffling terms in, 458–459
effective potential, 238–239; Coleman-Weinberg,
240; generated by quantum fluctuations, 243
Ehrenfest, P., 441Einstein, Albert: and cosmological constant, 449; and
repeated indices summation convention, 475n
Einstein-Hilbert action for gravity, 433–434;
Newtonian gravity derived from, 438; Yang-Millsaction compared with, 434–435
Einstein Lagrangian, linearized, 34Einstein’s theory of gravity, 81, 83; and deflection of
light, 439–440; gravitational waves in, 479–481;nonrenormalizability of, 172; Yang-Mills theorycompared with, 444–445, 513–520
electric charge: quantized, grand unification on,
410; quantum fluctuations and, 204, 205;renormalization of, 205
electric force: and gravitational force, comparison of,
29; quantum field theory on, 32–33
electromagnetic coupling, flow of, 358electromagnetic duality, 249electromagnetic field: Dirac equation in presence of,
101; as quantum field, 3–4; stress-energy tensorof, 83–84
electromagnetic force: between like charges, 33;
knowledge of, 448
electromagnetic wave, degrees of freedom of, 38electromagnetism: Faddeev-Popov method applied to,
185–186; and gravity, unification of, 442; Maxwellon (see Maxwell theory of electromagnetism);
weakness of, 414–415
electron(s): absolute identity of, 120–121, 134; binary
strings in, 428; Bose-Einstein statistics for, 120;in condensed matter system, 281; Cooper pairsof, 295; degrees of freedom of, 95, 99; effectof magnetic field on, 251–252; energy levelsavailable to, 5; Fermi-Dirac statistics for, 120;as fermions, 322; fractional Hall state of, 324;magnetic moment of ( seemagnetic moment
of electron); mass of, in classical physics, 180;noninteractive hopping, 298–299, 298f, 299f;pairing into bosons, 295; photon fluctuation into,200–202, 201f; photon scattering on, 152–157,153f, 157f; requirements of antisymmetric wave
function, 107; stability of, 413
electron-positron annihilation, 155–156, 389–391electron scattering, 132–143; cross sections for, 137–143; off electrons, 134–138, 135f; off nucleons,
deep inelastic, 386; off protons, 132–134, 133f,199; off protons, deep inelastic, 359; off protons,Schr ¨odinger equation for, 3; to order e
4, 145–149,
145–147f, 148f; potential, 133–134, 134f
electroweak theory, 170–171, 379; construction of,
379–383; renormalizability of, 384
energy: dark, 450; fundamental definition of, 83;
of mass, 35; quantum mechanics and specialrelativity on, 3; of vacuum (see vacuum energy)
energy density, 35energy-momentum tensor, 319energy scales: in particle physics, 169;
renormalization group and, 361
entropy, Boltzmann on, 311Euclidean path integral, 12
Euclidean quantum field theory, 287–288; and high-
temperature quantum statistical mechanics, 289;and quantum statistical mechanics, 289
Euler, Leonhard, 460Euler-Lagrange equation, 12, 80, 438, 448exact forms, 247
Fadeev-Popov method, 183–185, 267, 371; applying
to electromagnetism, 185–186; and derivation ofgraviton propagator, 437
Feng, Bo, 500Fermi, Enrico, 105, 137Fermi coupling, 170Fermi-Dirac statistics, 120Fermi field, mass correction to, divergence of, 180Fermi liquid, gapless modes in, 285fermion(s): and bosons, unification of, 461; degree
of divergence with, 178–179; electrons as, 322;Feynman rules for, 128–131, 128f; in lattice gaugetheory, 376; mass correction for, divergence of,180; massive Dirac, and Chern-Simons term,319–320
fermion-fermion scattering, Feynman diagram for,
172, 172f
fermion masses: in grand unification, 417–418;
naturally small, 419
fermion normalization factors, 134fermion propagator, 112Fermi theory of the weak interaction, 232; as effective
field theory, 456; intermediate vector boson and,171–172; nonrenormalizability of, 170, 179, 384;predictive power of, 453; within electroweaktheory, 171
ferromagnet(s), 229; effective low energy description
of, 344–345; low energy modes in, 345–346;
magnetic moments in, 344; order in, 328
ferromagnetic transition, 295Feynman, Richard: contribution of, 43; on difficulty
Index | 567
of quantum electrodynamics, 61; on Dirac, 105;
metaphorical language of, 113; on path integralformalism, 7–10; study of calculus by, 522; ontrace products of gamma matrices, 137; Yang-Millstheory and, 371
Feynman diagrams: beginning of, 29, 30f; breaking
shackles of, 311; canonical formalism and, 43;childish game generating, 53, 53f; connectedvs. disconnected, 29, 47; counterterms in, 175–176, 176f; Cutkosky cutting rule for, 215–216,217, 219; discovering, 43–51, 45f, 46f; dominanceof, 302; for electron scattering, 132–134, 133f,134f, 135f; evaluating, 538–539; for fermion-fermion scattering, 172, 172f; finite temperature,289; function of, 50; imaginary part of, 207–219,208f, 213f; limitations of, 67; loop, 45, 57–58,
57f, 58f, 181, 494; in momentum space, 54; newapproaches to, 483–486; orientation of, 54; pathintegral formalism and, 44; in perturbation theory,55, 56f; for photon scattering, 152, 153f, 155, 155f;regularization of, alternative ways of, 166; relatinginfinite sets of, 234–235; in spacetime, 54, 58, 213
Feynman gauge, 149Feynman rules, 534–537; colored, 485, 491, 495;
discovery of, 60; for fermions, 128–131, 128f; innonabelian gauge theory, 536–537; in physicalperturbation theory, 175–176, 176f; for quantumelectrodynamics, derivation of, 144–150; inrandom matrix theory, 397, 398f; for scalar field,54–55, 534–535; in spontaneously broken gaugetheories, 266–267; for vector field, 129, 130f,535–536; in Yang-Mills theory, 257, 257f, 494–495
field redefinition, 68–69, 218, 342field renormalization, 175field strength, construction of, 255–257Fierz, M., 121Fierz identities, 459Fisher, Matthew P. A., 336nFisher, Michael, 293; and renormalization groups,
361
fixed point(s), strong coupling, 359flux: fundamental unit of, 324; gauge potential and,
334
force: origin of, 29; particle and, 27–29. See also
specific force
forms: closed vs. exact, 247–248; geometric character
of, 250–251, 250f
fractional Hall effect, 323–324fractional (anyon) statistics, 315; coupling to
gauge potential, 316–317; gauge boson and, 320;misleading nature of term, 317; and quasiparticles,
327
freedom, degrees of, 37–38; canonical formalism
and, 67; Dirac equation and reduction in, 95; ofelectron, 95, 99; gauge invariance as redundancy
in, 268; longitudinal, in massive gauge field, 264;of photon, 186–187
free field theory (Gaussian theory), 21–23, 43; in
terms of Fourier transform, 26
Fujikawa, Kazuo, 278
gamma matrices, 94, 117, 538; products of, 95–
96; trace products of, evaluating, 136–137,153–154
Gamow, George, 120n“Gang of Four,” 366gapless mode, 284; Bogoliubov calculation of, 284;
linearly dispersing, 284–285
gauge boson(s): and fractional statistics, 320; and
intermediate vector boson, 309; mass spectrum
of, 266
gauge fixing, 183gauge invariance, 83n, 144, 475; of Chern-Simons
term, 328; and Dirac quantization of magneticcharge, 248; discovery of, 144n; in latticegauge theory, 376; in nonabelian gauge theory,preserving, 204; origin of, 183; proof of, 145–150,203–204; as redundancy in degrees of freedom,268; regularization respecting, 202–204; andrenormalizability, 411
gauge potential, 251; flux associated with, 334;
fractional statistics and, 316–317; in Hall fluid,325–326, 329; nonabelian, 254, 255
gauge theory(ies): Faddeev-Popov quantization of,
183–185, 267; and fiber bundles, correspondencebetween, 256; gravity, as, 436; lattice, 374–376;recent developments in, 497–512; redundancyin, 183–185, 189; S-matrix theory and, 498–
501; spontaneously broken, Feynman rulesfor, 266–267; spontaneously broken, magneticmonopoles in, 309; and superconductivitytheory, 296; symmetry breaking in, 263–265,268, 296; unsatisfactory formulation of, 474,497; vortex in, 307. See also nonabelian gauge
theory(ies)
gauge transformation (local transformation), 187,
254; and general coordinate transformation,443
Gauss-Bonnet theorem, 457, 459Gaussian integration, 14, 523Gaussian theory (free field theory), 21–23, 43; in
terms of Fourier transform, 26
Gell-Mann, Murray: and effective field theory, 460;
on quark color, 385; and seesaw mechanism, 426;σmodel of, 340–341; SU (3) of, 531; Yang-Mills
theory and, 371
Gell-Mann matrices, 265general coordinate invariance, 36
568 | Index
general coordinate transformation, and gauge
transformation, connection between, 443
general covariance, principle of, 81general relativity: finite size objects in, 480–482; and
quantum mechanics marriage of, 6; review of,84–86
generator, charge as, 80Georgi, Howard, 421n; grand unification theory of,
407–409
ghost fields, 372–374Giele, W ., 493Ginzburg, V ., 264; on London penetration length,
296; on second-order transitions, 292; onsuperconductivity, 295
Girvin, Steve, 328Glashow, Sheldon, 171; electroweak theory of, 383;
grand unification theory of, 407–409; and seesawmechanism, 426; Yang-Mills theory and, 379
gluon(s), 386; origins of concept, 235gluon scattering, 483–496; approaches to calculation
of, 483–493, 484f, 491f; spinor helicity formalismand, 486–491, 496
Goldberger, Murph L., 105, 137, 460Goldberger-Treiman relation, 235, 342Goldstone’s theorem, 228–229; in condensed matter
physics, 229
Golfand, Yu. A., 461Gordon decomposition, 195Goto, T ., 469grand unification, 452; binary code in, 426–
428; charge conjugation in, 429; and deeperunderstanding of physics, 410–411; fermionmasses in, 417–418; and freedom from anomaly,411–412, 429; and hierarchy problem, 419;need for, 407; and origin of matter, explanationfor, 418; and proton decay, 413–414, 415,456; SO (10): antineutrino field in, 425–426;
SO (18), 428; spinor representation of, 421–
423, 424, 426; SU (5), 531; SU (5), Georgi
and Glashow theory of, 407–409; triumph of,415–416
Grant, A. K., 516Grassmannian symmetry, 355Grassmann integration, 126–127Grassmann number(s), 123, 126; in path integral for
spinor field, 124
Grassmann variables, 246gravitational force. See gravity
gravitational interaction, 36gravitational waves, and effective field theory,
479–482
gravition propagator, 437–438graviton: coupling to matter, 435; definition of, 83;
deformed polarizations of, 514–515; as elementaryparticle, 434, 448; force associated with, 29; in
(n+3+1)-dimensional universe, 41–42; recent
developments on, 513–520; self-interaction of,434; in spacetime, 515–517; spin of, 35, 39; instring theory, 513–514, 516
gravity: Einstein-Hilbert action for, 433–434;
Einstein on ( seeEinstein’s theory of gravity);
and electromagnetism, unification of, 442;as field theory, 434–436; as gauge theory,436; helicity structure of, 446; of light, 441;Newton on (see Newton’s gravitational force);
nonrenormalizability of, 434; weak field actionfor, 436–437
Green’s function(s), 47, 55, 352; generating, 50;
propagator related to, 23
Gross, David, 386
Gross-Neveu model, 402–404ground state, in quantum field theory, 37, 225ground state degeneracy, 319group theory, review of, 525–533. See also special
orthogonal group SO (N); special unitary group
SU (N)
hadron(s): electron-positron annihilation into, 389–
391; in electroweak unification, 383; experimentalobservation of, 231; quarks as components of, 385
Hall effect, 351–352; fractional, 323–324; integer, 323Hall fluid(s), 322–330; Chern-Simons term for,
326; effective field theory of, 324–325, 452–453; electron tunneling in, 329; five generalstatements/principles of, 325, 329; gauge potentialin, 325–326, 329; incompressibility of, 323, 328;Laughlin odd-denominator, 327; order in, 328
handedness, field, 100; charge conjugation and, 101Hansson, T ., 324harmonic paradigm, 5Hasslacher, Brosl, 405Hawking radiation, 290–291Heaviside, O., 24, 245hedgehog, 308Heisenberg, Werner: approach to quantum
mechanics, 61–62; and effective field theory, 460;isospin SU (2) of, 531; isospin symmetry of, 387,
388; on neutron and proton, symmetry of, 77
helicity, topological quantization of, 532–533helicity formalism, spinor, 486–491, 496, 501, 521hierarchy problem, grand unification and, 419Higgs field, covariant derivative of, 266Higgs particle, mass of, 384high energy physics: renormalization group in,
359–360
high frequency behavior, 208–210Hofstadter, R., 199homotopy groups, 307
Index | 569
Hopf term, 318; non-local, 329
Howe, P., 470Hubbard-Stratonovich transformation, 192
identity, absolute, 120–121, 134
imaginary part, of Feynman diagrams, 207–219,
208f, 213f
impurities, 323; condensed matter physics and study
of, 354; and random potential, 350
infinities, in quantum field theory, 161–162instanton(s), 309; discovery of, 473integer Hall effect, 323integration measure, in path integral formalism, 67integration variables, shifting, 272interchange symmetry, 77internal symmetry, 77
inverse square law, 40irrelevant operators, 363Ising model, 361isospin symmetry of Heisenberg, 387, 388Itzykson, Claude, 402Iwasaki, Y ., 439
Johansson, H., 519
Jona-Lasinio, G., 237Jordan, P., 107Josephson junction, fundamental relation
underlying, 192
Kadanoff, Leo, 361; and renormalization groups,
361
Kaluza-Klein compactification, 442–443; derivation
of, 447
Kardar-Parisi-Zhang equation, 347Kawai, H., 513kinks. See solitons
Kivelson, Steve, 324Klein-Gordon equation, 21, 93, 95, 190; Schr ¨odinger
equation derived from, 190
Klein-Gordon operator, 113Klein-Nishina formula, 155Kockel, B., 460Kosterlitz-Thouless transition, 310Kramer’s degeneracy, 103
Lagrangian: Dirac (see Dirac equation); gauge
invariant, 253–254; Maxwell (see Maxwell
Lagrangian); Meissner, 332, 335; as mnemonic,340, 342; for quantum electrodynamics, 101, 144;symmetries of breaking, 223; weak interaction,100; Yang-Mills, 257
Lamb shift in atomic spectroscopy, 205Landau, L. D., 264; on complex momenta, 498; on
London penetration length, 296; on second-ordertransitions, 292; on superconductivity, 295; on
superfluidity, 284
Landau gauge, 149Landau-Ginzburg approach to quantum field theory,
18
Landau-Ginzburg theory (mean field theory),
292–294; order in, 328
Landau levels, 323Laplace, P.-S., 290Large Hadron Collider, 483large Nexpansion, 394–396; Dyson gas approach to,
400–402; field theories in, 402–404
Larmor circle, 322, 323lattice gauge theory, 374–376; Wilson loop in,
376–377, 457
Laughlin odd-denominator Hall fluids, 327
Lee, B., 173Lee, D. H., 336nLee, T sung-dao, 100Legendre transform, 238–239Leinaas, J., 315length scales: in condensed matter physics, 169;
renormalization group and, 361–362
leptons: families of, 384; generations of, 428; and
quarks, neutral current interaction between, 383
L´evy, M., σmodel of, 340–341
Lewellen, David C., 513Licciardello, D., 366light, gravity of, 441light beam, stress-energy tensor of, 445Likhtman, E. P., 461linearly dispersing mode, 284; velocity of, 285local field theory, 474, 521–522localization: Anderson, 351, 354; Anderson, in
renormalization group language, 366–367; studyof, 355
local transformation (see gauge transformation)
logarithmic divergence, 175, 176–177London penetration length, 296loop diagrams, 45, 57–58, 57f, 58f, 181, 494Lorentz algebra, 114–116Lorentz boosts, 114–115Lorentz group: defining representation of, 116;
generators of, algebra for, 115–116; spinorrepresentation of, 116–118
Lorentz invariance, 475; canonical formalism and,
63, 66–67; Euclidean equivalent of, 362; inquantum field theory, 18, 24; recent developmentsin, 507–510
Lorentz transformation: and Dirac equation, 96–97Low, F., 460
L¨uders, G., 121
MacDonald, Alan, 328
570 | Index
magnetic charge (monopole), 249, 309; confinement
in superconductor, 386–387; Dirac quantizationof, 248–249, 252; duality of, 334; electricallycharged (dyon), 309; mass of, 309; and Maxwell’sequations, 249; quantum mechanics and, 245; inspontaneously broken gauge theory, 309
magnetic moment of electron: anomaly in, 196;
calculation of, 454; in Dirac equation, 194–195;Schwinger on, 196–198, 454
magnetic moment of ferromagnet and
antiferromagnet, 344
magnetic moment of proton, anomaly in, 454–455Majorana, Ettore, 102n, 543Majorana equation, 102Majorana mass, 102; for neutrino, 102Majorana spinor, 102, 543
Mandelstam variables, 137–138, 498marginal operators, 364mass(es): attraction between, 35–36; of electron,
180; energy of, 35; of gauge boson, 266; of Higgsparticle, 384; of magnetic charge (monopole), 309;Majorana, 102; of neutrino, 102; of nucleon, 341;Planck, 41–42, 434; of soliton (kink), 305
massive gauge field, Nambu-Goldstone boson and,
264–265
massive spin 1 field, vs. massless spin 1 field, 183massive spin 1 particle: degrees of freedom of, 38;
degrees of polarization of, 34; field theory of,32–33; propagator for, 34; in Yang-Mills theory,379
massive spin 2 particle: degrees of polarization of,
35; propagator for, 35, 439
mass renormalization, 174matrix (matrices): gamma ( seegamma matrices);
Gell-Mann, 265; Pauli, 265; without inverse,182–183
matter: dark, 450; origin of, explanation for, 418;
states of, 328
mattress model of scalar field theory, 4–5, 4f;
disturbing, 21, 21f; path integral description of,17–19
Maxwell, James Clerk, 521Maxwell action, 182Maxwell equations, magnetic charges and, 249Maxwell Lagrangian, 32, 34, 84; bypassing, 33–34;
derivation of, 38
Maxwell term, 320, 329Maxwell theory of electromagnetism: development
of, 474–476; Yang-Mills theory compared with,257
mean field theory (Landau-Ginzburg theory),
293–294; order in, 328
Meissner effect, 296, 386Meissner Lagrangian, 332, 335meson(s): birth of, quantum field theory on, 55–56,
56f;π(see pion[s]); σ, 341, 342; soliton compared
with, 304; vector, field theory of, 32–33. See also
massive spin 1 particle
meson-meson scattering amplitude, 357; canonical
formalism and, 64–65; cutoff dependence of,173; dimensional analysis on, 173; divergenceof, 161–162; path integral formulation of, 166;regularization and, 163; renormalization and,164–166
Michell, John, 290Mills, Robert, and nonabelian gauge theory, 253,
255
Minkowski, Peter, and seesaw mechanism, 426Minkowskian path integral, 287Minkowskian spacetime, 36
momentum: complex, 498–501, 499f; fundamental
definition of, 83; orbital angular, Dirac equationon, 194–195; spin angular, Dirac equation on, 195;square root of, 486–489
momentum density, in nonrelativistic theory, 191momentum space, 26; fermion propagator in, 113;
Feynman diagrams in, 54
monopole. See magnetic charge
Montonen, J., 334muon, weak decay of, 380Myrheim, J., 315
Nambu, Yoichiro, 297, 469; Nobel prize for, 228n
Nambu-Goldstone boson(s), 228–229; gapless mode
as, 284; in massive gauge field, 264–265; π
mesons (pions) as, 234, 387, 388; in relativistic vs.nonrelativistic theories, 285
naturalness, notion of, 419N´eel state, for antiferromagnet, 346
Ne’eman, Y ., SU (3) of, 531
neutral current interaction, 383neutrino(s): handedness of, 101; mass of, 102neutrino masses, effective field theory of, 456neutron(s): βdecay of, 231–232; electric dipole
moment for, 259; and proton, internal symmetryof, 77
Neveu, Andr ´e, 402, 405
Newton’s gravitational force: and Coulomb’s electric
force, comparison of, 29; derived from Einstein-Hilbert action, 438; quantum field theory on, 32,33–36
Noether current, 191, 234Noether’s theorem, 78–79, 100, 341; elaborate
formulation of, 80
nonabelian Berry’s phase, 261, 346
nonabelian gauge potential, 254, 255; coupling to a
fermion field, 260
nonabelian gauge theory(ies), 253–260; chiral
Index | 571
anomaly in, 276; differential forms in, 255,
256; Feynman rules in, 536–537; gaugeinvariance in, preserving, 204; ghost actionin, 372; redundancy of, Faddeev-Popovapproach to, 183; renormalizability of, 173,411; strong interaction described by, 259, 379;’t Hooft double-line formalism and, 258–259;unsatisfactory features of, 474. See also Yang-Mills
theory
nonanalyticity: emergence of, 292; symmetry
breaking and, 293
noncommutative field theory, 474nonrenormalizable theory(ies), 169, 179, 453;
counterterms in, 179, 241; Einstein’s theoryof gravity as, 172; Fermi’s theory of the weakinteraction, 170, 179
nonrenormalization of the anomaly, 277notation, dotted and undotted, 116–117, 541–544;
replacing, 475
nucleon(s): attraction between, 28; electron scattering
off of, deep inelastic, 386; mass of, 341; and pions,interaction between, 340–341; wave function ofquarks in, 385
Olive, D. I., 334
optical theorem, 215–216, 219orbital angular momentum, Dirac equation on,
194–195
order parameters, 295orthogonal groups, embedding unitary groups into,
423–424
Parisi, Giorgio, 402
parity, 98; Dirac equation and, 98; and Dirac spinor,
117; weak interaction and, 100, 379–380
Parke, S. J., 493particle(s): birth and death of, 4–5; birth of, quantum
field theory on, 55–56; field associated with, 26–27; force associated with, 27–29; interchanging,315–316, 316f; propagation of, describing, 48–49, 50; scattering of (see scattering of particles);
sources and sinks for, 20. See also specific particles
particle physics: and condensed matter physics, 281,
452–453; energy scales in, 169; family problem in,428; spontaneous symmetry breaking in, 292, 297,449
partition function, in quantum statistical mechanics,
288–289
path integral formalism: vs. canonical formalism,
44, 61, 67; chiral anomaly and, 278; and classicallimit, 19; derivation of, 44; description of mattress
model, 17–19; Dirac on, 10–13; Feynman on, 7–10; Grassmann math and, 127; history of, 60;integration measure in, 67; replacing, 475; forspinor field, 124; and vacuum energy, calculation
of, 123–125
Pauli, Wolfgang, on spin-statistics connection, 121Pauli exclusion principle, 120, 323; history of, 120nPauli-Hopf identity, 345Pauli matrices, 265Pauli-Villars regularization, 75, 166–168Peierls, Rudolf, 365Peierls instability, 300pentagon anomaly, 276, 277fperturbation theory, 49–51; bare, 175; Feynman
diagrams in, 55, 56f; finite temperature, 289;physical (renormalized/dressed), 175–176, 176f
perturbative quantum gravity, 441ϕ
4theory, renormalizability of, 173, 175
phonon(s), 5, 284
photon(s): absence of rest frame for, 186–189; birth
and death of, 4; Bose-Einstein statistics for, 120;degrees of freedom of, 186–187; electron-positronannihilation into, 155; emission and absorption of,150; fluctuation into electron and positron, 200–202, 201f; force associated with, 29; longitudinalmode of, 150; spin of, 36, 39
photon propagation: charge as measure of, 204;
quantum fluctuations and, 200–202, 201f
photon propagator, 149–150; Fourier transform of,
205; physical (renormalized), 201, 201f
photon scattering, 152–157; cross sections for,
152–155; on electrons, 152–157, 153f, 157f
physical perturbation theory, 175–176, 176fpion(s) (π meson): massless, 235, 341; Nambu-
Goldstone boson, 388; as Nambu-Goldstoneboson, 234, 387; and nucleons, interactionbetween, 340–341; prediction regarding, 29;quarks as components of, 385; weak decay of,231–233
pion-nucleon coupling constant, 235Planck mass: modified, 434; for (n+3+1)-
dimensional universe, 41–42
Planck’s constant, 181Podolsky, B., 441Poincar ´e lemma, 247
point particle: action of, constructing, 84–86; stress
energy of, calculating, 86; world line traced out by,length of, 84, 85f
Poisson equation, 438polarization, degrees of, 34Politzer, H. D., on Yang-Mills theory, 386Polyakov, Alexander (Sasha), 498; on magnetic
monopoles, 309
Polyakov action, 470
Pontryagin index, 310positron(s): Dirac’s conception of, 5; photon
fluctuation into, 200–202, 201f
572 | Index
potential energy, double-well, 224, 224f
power counting theorem, 176–178preons, theories about, 278product rule, 445propagation of particles, describing, 48–49, 50propagator, 23–25; in canonical formalism, 67–68;
for Dirac field, 127; fermion, 112; graviton, 437–438; for massive spin 1 particle, 34; for massivespin 2 particle, 35, 439; photon, 149–150
proton(s): charge of, grand unification on, 410;
electron scattering off of, 132–134, 133f, 199;electron scattering off of, deep inelastic, 359;electron scattering off of, Schr ¨odinger equation
for, 3; magnetic moment of, anomaly in, 454–455;and neutron, internal symmetry of, 77; quarks ascomponents of, 385; stability of, 413
proton decay: branching ratios for, 416–417; effective
theory of, 455–457; grand unification and,413–414, 415, 456; slow rate of, 418
quantum chromodynamics (QCD), 360, 386; analytic
solution of, search for, 391; at high energies, 391;largeNexpansion of, 394–396; renormalization
group flow of, 388–389
quantum electrodynamics (QED), 32; coupling
constant of, 164; coupling in, 358; electromagneticgauge transformation in, 189; Feynman ondifficulty of, 61; Feynman rules for, derivationof, 144–150; intellectual incompleteness of, 121;Lagrangian for, 101, 144; renormalizability of,173
quantum field theory(ies): in (0+0)-dimensional
spacetime, 397; in 2-dimensional spacetime,470; anharmonicity in, 43, 89; asymptoticbehavior of, study of, 359–360; central identityof, 182, 523; and condensed matter physics,5, 190, 281; crisis of, 231, 340, 452; in curvedspacetime, 82, 290; divergences in, 57–58, 161–162; Euclidean, 287–288, 289, 290; at finitedensity, 291; at finite temperature, 289–290;gravity as, 434–436; ground state in, 37, 225;harmonic paradigm and, 5; hidden structuresin, 476; history of, 60; infinities in, 161–162;innovative applications of, 473–474, 476; integralof, 88–89; low energy manifestation of, 162, 169,452; mattress model and, 17–19; motivationfor constructing, 55; need for, 3–5, 6, 123;nonrelativistic limit of, 190–191; relativisticvs. nonrelativistic, 191–193; renormalizablevs. nonrenormalizable, 169; on repulsion andattraction, 32–36; restrictions within, 474; steps
toward, 235; strong and weak interactionsapplied to, 231; of strong interaction, 340;supersymmetric, 461, 467–468; surface growthand, 347–349; symmetry breaking in, 225–226;theories subsumed by, 473; threshold of ignorance
in, 162–163, 453; triumph of, 452, 473; vacuumin, 20
quantum fluctuations: axial current conservation
destroyed by, 274–275; effective potentialgenerated by, 243; and electric charge, 204, 205;first order in, 239–240; higher order, and chiralanomaly, 310; and photon propagation, 200–202,201f; and symmetry breaking, 229, 237, 242, 270
quantum Hall fluid. See Hall fluid(s)
quantum Hall system, 281quantum mechanics: antimatter as requirement
in, 157; and general relativity, marriage of, 6;harmonic oscillator in, solving, 43; Heisenberg’sapproach to, 61–62; and magnetic monopoles,245; partition function in, 288–289; path integral
formalism of, 7–12; quantum field theory asgeneralization of, 88–89, 473; and relativisticphysics, joining in spin-statistics connection,122; and special relativity, marriage of, 3, 6,121; symmetry breaking in, 225–226; symmetryof, 270; time reversal in, 102–104; and vectorpotential, need for, 245
quantum statistics, 120quantum vacuum, 358quark(s): color of, 385, 386; confinement of, 377,
386–387; in electroweak unification, 383; familiesof, 384; flavors of, 385; generations of, 428; andleptons, neutral current interaction between,383; origins of concept, 235; strong interactionbetween, weakening of, 360
quasiparticle(s), 326; charge of, 327; fractional
statistics and, 327; as vortex, 328
radiation: and atoms, interaction between, 3;
Hawking radiation, 290–291
Ramakrishnan, T . V ., 366Ramond, Pierre, and seesaw mechanism, 426random dynamics, and quantum physics, 349random matrix theory, 396–397; Feynman rules in,
397, 398f
random potential, impurities and, 350Rarita-Schwinger equations, 119Rayleigh, Lord, 458recursion, 501–503, 507–512, 521; BCFW, 500, 507,
514
redundancy, Faddeev-Popov approach to, 183–185reflection symmetry, 76, 226; breaking, 223, 224,
225
Regge, T ., 498regularization, 163; Casimir force and, 71–75;
dimensional, 167, 168, 204; gauge invariancerespected by, 202–204; Pauli-Villars, 75, 166–167
relativistic physics: equations of motion in, unified
view of, 95; language of, 26; and quantum physics,
Index | 573
joining in spin-statistics connection, 122. See also
general relativity; special relativity
relativistic quantum field theory: correctness of,
establishment of, 196; vs. nonrelativistic quantumfield theory, 191–193
relevant operators, 363renormalizable conditions, imposing, 241–242renormalizable theory(ies), 169, 173, 453; electroweak
theory as, 384; nonabelian gauge theory as, 173,411;ϕ
4theory as, 173, 175, 178–179; Yukawa
theory as, 179
renormalization, 161, 164–166; coupling, 173–174;
of electric charge, 205; field, 175; mass, 174; wavefunction, 175
renormalization group, 356, 358; and Anderson
localization, 366–367; in condensed matter
physics, 360–363; and effective description,367; effective field theory philosophy and, 453;in high energy physics, 359–360; in quantumchromodynamics, 388–389
renormalization theory, application of, 240–241renormalized coupling constant, 166renormalized (dressed) perturbation theory, 175–176,
176f
reparametrization invariance, 84replica method, 353–354representations: conventions for naming, 526;
multiplying, 530–531
repulsion: of bosons, 192–193, 283, 337; quantum
field theory on, 32–33; spin 1 particle and, 36–37;of vortices, 338
rest frames, for photons, absence of, 186–189Ricci tensor, 433Riemann-Christoffel symbol, 85, 445Riemann curvature tensor, 433, 480–481, 515–516Riemannian manifolds, differential geometry of,
443–444
Rosenbluth, Marshall, 105rotation group, 114; and Lorentz group, symmetry
of, 118
R
ξgauge, 267, 268
Salam, Abdus, 171; electroweak theory of, 383;
superspace and superfield formalism of, 462, 463
scalar boson operator, 113scalar field: complex, 65–66; Feynman rules for, 54–
55, 534–535; quantizing in curved spacetime, 82;and vacuum energy, 66
scalar field theory: classical field equation in, 20;
Euclidean functional integral and, 287; Euclideanversion of, 293; massless version of, 284; simplicity
of, 519–520
scalar potentials, 245scattering of particles: describing, 51–53, 51f, 52f;
fermion-fermion, Feynman diagram for, 172;meson-meson (see meson-meson scattering
amplitude); reflection symmetry in, 76; andvacuum fluctuations, 124. See also electron
scattering
Schouten identity, 492, 493Schrieffer, Bob, 297Schr ¨odinger equation: electromagnetic gauge
transformation in, 189; Klein-Gordon equationand, 21n, 190; limitations of, 3; Yang-Millsstructure in, 261
Schwarz, John, on string theory, 470nSchwarzschild black hole, 311Schwarzschild solution, for Hawking radiation, 290Schwinger, Julian: on complex plane, 208;
and effective potential, 237; on Feynman’scontribution, 43, 50, 56; on magnetic moment
of electron, 196–198, 454; on path integralformalism, 60; at Pocono conference (1948), 105;teaching style of, 454; Yang-Mills theory and,379
second-order phase transitions, 292seesaw mechanism, 37, 426Seiberg, Nathan, 334self-dual theory, 337semions, 315σmeson, 341, 342
σmodel, 340–341; for ferromagnets and
antiferromagnets, 345, 346; nonlinear, 342, 346
sky color, effective field theory of, 457–458Slansky, Dick, and seesaw mechanism, 426S-matrix theory, 68, 235, 340, 498–501
solid state physics, Dirac equation in, 298, 299solitons (kinks): discovery of, 302–304, 473;
dynamically generated, 400–405; mass of, 304;topological stability of, 304; unifying language fordiscussing, 307
SO (N).See special orthogonal group
sources and sinks, creating, 20, 51, 51fspacetime: curved (see curved spacetime); dimension
of, and symmetry breaking, 229; discretizing, 22;Feynman diagrams in, 54, 58, 213; gravitationalwaves in, 479–482; graviton in, 515–517; symmetryof, Lorentz invariance as, 76
special orthogonal group SO(N), 525–527; binary
code in, 427–428; review of, 531–532; SO(3), 526–
527; SO(10) grand unification, antineutrino field
in, 425–426; SO(18), 428; spinor representation
of, 421–423, 424, 426
special relativity: antimatter as requirement in, 157;
and quantum mechanics, marriage of, 3, 6, 121
special unitary group SU (N), 527–530; decomposing
representations of, 531; of Heisenberg, 531; SU
(2), 529–530; SU (3), 529, 530; SU (3), of Gell-
Mann and Ne’eman, 531; SU (5), 531; SU (5),
Georgi and Glashow theory of, 407–409
574 | Index
spin angular momentum, Dirac equation on, 195
spinor(s): Dirac, 94, 96, 114, 117; Majorana, 102;
representations of, 116–117; Weyl, 117, 462
spinor field: deriving, 125–127; path integral for, 123;
path integral for, Grassmann numbers in, 124;vacuum energy of, 111–112, 125
spinor helicity formalism, 486–491, 496, 501, 521spin-statistics rule, 120–121; and anticommutation
relations, 122, 123; price of violating, 121–122
spin wave, 229spontaneous symmetry breaking, 224, 225, 227;
continuous: and massless fields, 228–229; ingauge theories, 263–265; in particle physics,292, 297, 449; quantum fluctuations and, 229;of reflection symmetry, 225; in relativistic vs.nonrelativistic theories, 285; second-order phase
transitions and, 292; and superfluidity, 283–284
square anomaly, 276, 277fsquare root of momentum, 486–489steepest-descent approximation, 16Stokes’ theorem, 307Stoner, E. C., 120nStrathdee, J., superspace and superfield formalism
of, 462, 463
stress-energy tensor, 35; definition of, 83; of light
beam, 445; properties of, 84
string theory: 2-dimensional field theory, 469–
470; as candidate for unified theory, 433, 452;and cosmological constant problem, inabilityto resolve, 450; duality of, 334; future of, 513;graviton in, 515–517; Kaluza-Klein idea and, 442;origins of, 6, 387; p-forms in, 251; in quantum
field theory, 473; Schwarz on, 470n
strong coupling: fixed point in, 359; linking to
perturbative weak coupling, 473
strong interaction: chiral symmetry of, 234; currently
accepted theory of, 379; fundamental theoryof, 360; hadronic, 36; at low energies, 340–341;nonabelian gauge theory on, 259, 379; quantumfield theory of, 235, 340; renormalization groupflow applied to, 368; symmetries of, 234, 387–388
SU (N).See special unitary group
supercharges, 464superconductivity, 295–297; and Meissner effect, 296superconductor(s): monopole confinement in,
386–387; type II, flux tube in, 307
superfield, 464–465; chiral, 464, 466; vector, 466–467superfluidity, 192; gapless excitations and, 284–285;
Lagrangian summarizing, 284; linearly dispersingmode of, 284; spontaneous symmetry breakingand, 283–284
superspace and superfield formalism, 462–463superstring theory, 470supersymmetric action, 466supersymmetric algebra, 462–463
supersymmetric field theories, 461, 467–468;
Yang-Mills, 392, 467–468
supersymmetric method, 355supersymmetric transformation, total divergence
under, 465–466
supersymmetry, 112; Dirac spinor and, 114;
inventing, 462; motivations for, 461
surface growth, 347, 360; and quantum field theory,
348–349
Swieca, Jorge, 230nsymmetry, 76–80; in amplitudes, 78; breaking,
226; chiral, 234, 387, 388, 419; classical vs.quantum, 270–271; conserved current and, 78–79; continuous, 77–78, 226; in field theories,475; Grassmannian, 355; Heisenberg isospin,
387, 388; interchange, 77; internal, 77; powerof, 18, 76, 118; reflection, 76, 226; replica, 353;of spacetime, Lorentz invariance as, 76; stronginteraction, 234, 387–388; tensors and, 526, 528.See also supersymmetry
symmetry breaking, 223–230; continuous symmetry
and, 226; dimension of spacetime and, 229;dynamical, 230, 388; in gauge theories, 263–265, 268, 296; and nonanalyticity, 293; quantumfluctuations and, 229, 237, 242, 270; in quantummechanics vs. quantum field theory, 225–226;reflection symmetry and, 223, 224, 225; andsuperfluidity, 283–284; and vacuum energy, 449.See also spontaneous symmetry breaking
Taylor, T . R., 493
Teller, Edward, 105temperature: black hole, 290; and cyclic imaginary
time, 289; finite, quantum field theory at, 289–290
tensor(s): energy-momentum, 319; of light beam,
445; of orthogonal group, 525–526; Ricci, 433;Riemann curvature, 433; stress-energy, 35, 83–84; symmetry properties of, 526, 528; of unitarygroup, 527–528; vacuum polarization, 200, 201f,204, 208, 209f, 211, 216, 218
tensor field, 35, 83θterm, 259
Thomas precession, 115’t Hooft, Gerardus, 173; on electroweak theory,
384; on large Nexpansion, 394; on magnetic
monopoles, 309
’t Hooft double-line formalism, 258–2593-brane, 40–42time ordering in canonical formalism, 67–68time reversal, 102–104; and Dirac equation, 104
Tolman, R., 441Tolman-Ehrenfest-Podolsky effect, 441, 446Tomonaga, Shin-Itiro, 60
Index | 575
topological current, 304
topological field theory, 318topological objects, 306; discovery of, shock of, 311.
See also specific objects
topological order, 328topological quantum fluids, 322. See also Hall fluid
total divergence, under supersymmetric
transformation, 465–466
trace, 526tree diagrams, 45, 483–484, 491–494Treiman, Sam, 236. See also Goldberger-Treiman
relation
triangle anomaly, 271f, 276twistor space, 494–495Tye, Henry, 513
ultraviolet catastrophe, 448
ultraviolet divergence, 162uncertainty principle, 3, 290unification. See grand unification
unitarity, 215–216, 500unitary gauge, 267unitary groups, embedding into orthogonal groups,
423–424
universe: 3-brane, 40–42; early, 290; formation of
structure in, 36
vacuum: disturbing of, 20, 21f, 70–73; quantum, 20,
358
vacuum energy: calculation of, using path integral
fomalism, 123–125; disturbance of vacuum and,70–73; of free scalar field, 66; of free spinor field(Dirac field), 111–112, 125; Grassmann pathintegral for, 127; symmetry breaking and, 449
vacuum expectation value, 226vacuum fluctuations, 59–60, 59f; Feynmann diagram
corresponding to, 129; scattering of particles and,123
vacuum polarization tensor, 200, 201, 204, 208, 209f,
211, 216, 218
van Dam, H., 439van der Waerden notation. See dotted and undotted
notation
vector field, interacting with Dirac field, 100;
Feynman rules for, 129, 129f, 535–536
vector meson (massive spin 1 meson): field theory
of, 32–33. See also massive spin 1 particle
vector potential, 245vector superfield, 466–467Veltman, Tini, 173, 439; on electroweak theory, 384vielbeins, 443
visual perception, application of field theory to, 476vortex (vortices), 306, 331–332; as charges in dual
theory, 332–334; density of, 333; duality of, 334;as flux tube, 307; motion in fluid, 338–339, 338f;
paired with antivortex, 310–311, 339; quasiparticleas, 328; repulsion of, 337
Ward-Takahashi identity, 149, 411
wave function(s), Anderson localization of, 351, 354wave function renormalization, 175wave packets, in mattress model, 4–5, 4fweak interaction, 37; intermediate vector boson of,
171–172, 309; and parity, 100, 379–380; quantumfield theory applied to, 231. See also Fermi theory
of the weak interaction
weak interaction Lagrangian, 100Weinberg, Steve, 171, 508; electroweak theory of, 383Weisskopf phenomenon, 180–181; grand unification
and, 419
Wen, Xiao-gang, 324, 328; and topological order, 328Wentzel, Gregory, 105Wess-Zumino model, 462Weyl basis, 98–99, 118Weyl-Eddington terms, 457Weyl spinors, 117; and supersymmetry, 462Wheeler, John, 365nWick, Gian Carlo, 14Wick contractions, 14–16, 47Wick rotation, 12, 287Wick theorem, 14Wigner, Eugene: on antisymmetric wave function
of electron, 107; and law of baryon numberconservation, 413; and random matrix theory, 396;on time reversal, 102
Wigner semicircle law, 397–400Wilczek, Frank, 315, 316; on Yang-Mills theory, 386Wilson, Ken, 161; and complete theory of critical
phenomena, 293; and effective field theoryapproach, 452; and lattice gauge theory, 374–376;and renormalization groups, 361
Wilson loop, 261; in lattice gauge theory, 376–377,
457; and quark confinement, 386
Witten, Ed, 334, 500Wu, Tai-tsun, 248
Yanagida, T ., and seesaw mechanism, 426
Yang, Chen-Ning, 100, 105, 248; and nonabelian
gauge theory, 253, 255
Yang-Mills bosons, 257, 386; self-interaction of,
434
Yang-Mills coupling constant, 258–259Yang-Mills Lagrangian, 257Yang-Mills theory, 257–258; area law in, 377;
asymptotically free, 386; Einstein-Hilbert action
compared with, 434–435; Einstein’s theoryof gravity compared with, 444–445, 513–520;Feynman rules in, 257, 257f, 494–495;
576 | Index
Yang-Mills theory (continued)
gluon scattering in, 483–496; original responseto, 371, 379; perturbative approach to, 374;quantizing, 371–373; recent developmentsin, 483–496, 501–504, 513–520; recursion in,501–503; Schr ¨odinger equation and, 260–
261; supersymmetric, 392, 467–468; Wilsonformulation of, 374–376
Young tableaux, 526Yukawa, H., 28–29, 171Yukawa coupling, 170
Yukawa theory, renormalizability of, 178–
179
Zakharov, V ., 439
Zee, A., 316Zhang, Shou-cheng, 324Zinn-Justin, Jean, 173Zuber, Jean-Bernard, 402Zumino, Bruno, 121, 470