Topics_in_particle_physics_and_cosmology
PDF · 124 pages · 1.0 MB
Open PDF file
A doctoral thesis by Alejandro Jenkins, advised by Mark Wise, defended at Caltech in May 2006 (arXiv hep-th/0607239). It covers massless spin-1 and spin-2 particles, gauge and diffeomorphism invariance, the Weinberg-Witten theorem, Goldstone photons and gravitons, phenomenology of spontaneous Lorentz violation, dark energy and anthropic arguments, and the reverse sprinkler problem. It is someone else's work held in the archive, not by Phil.
AI-written summary; may contain errors.
Extracted text (machine-read; may contain errors)
arXiv:hep-th/0607239 v1 29 Jul 2006Topics in Theoretical Particle Physics and
Cosmology Beyond the Standard Model
Thesis by
Alejandro Jenkins
In Partial Fulfillment of the Requirements
for the Degree of
Doctor of Philosophy
California Institute of Technology
Pasadena, California
2006
(Defended 26 May, 2006)
ii
c∝ci∇cleco√y∇t2006
Alejandro Jenkins
All Rights Reserved
iii
Para Viviana, quien, sin nada de esto, hubiera sido posible ( sic)
iv
If I have not seen as far as others, it is because giants were st anding on my shoulders.
— Prof. Hal Abelson, MIT
v
Acknowledgments
How small the cosmos (a kangaroo’s pouch would hold it), how p altry and
puny in comparison to human consciousness, to a single indiv idual
recollection, and its expression in words!
— Vladimir V. Nabokov, Speak, Memory
What men are poets who can speak of Jupiter if he were like a man , but if
he is an immense spinning sphere of methane and ammonia must b e silent?
— Richard P. Feynman, “The Relation of Physics to Other Scien ces”
I thank Mark Wise, my advisor, for teaching me quantum field th eory, as well as a great
deal about physics in general and about the professional pra ctice of theoretical physics. I
have been honored to have been Mark’s student and collaborat or, and I only regret that,
on account of my own limitations, I don’t have more to show for it. I thank him also for
many free dinners with the Monday seminar speakers, for his p atience, and for his sense of
humor.
I thank Steve Hsu, my collaborator, who, during his visit to C altech in 2004, took me
under his wing and from whom I learned much cosmology (and wit h whom I had interesting
conversations about both the physics and the business world s).
I thank Michael Graesser, my other collaborator, with whom I have had many oppor-
tunities to talk about physics (and, among other things, abo ut the intelligence of corvids)
and whose extraordinary patience and gentlemanliness made it relatively painless to expose
to him my confusion on many subjects.
I accuse D´ onal O’Connell of innumerable discussions about physics and about such
topics as teleological suspension, Japanese ritual suicid e, and the difference between white
wine and red. Also, of reading and commenting on the draft of C hapter 2, and of quackery.
I thank Kris Sigurdson for many similarly interesting discu ssions, both professional and
unprofessional, for setting a ridiculously high standard o f success for the members of our
vi
class, and for his kind and immensely enjoyable invitation t o visit him at the IAS.
I thank Disa El ´iasd´ ottir for many pleasant social occasions and for loudl y and colorfully
supporting the Costa Rican national team during the 2002 Wor ld Cup. ¡Ticos, ticos! I
also apologize to her again for the unfortunate beer spillin g incident when I visited her in
Copenhagen last summer.
I thank my officemate, Matt Dorsten, for patiently putting up w ith my outspoken fond-
ness for animals in human roles, for clearing up my confusion about a point of physics on
countless occasions, and for repeated assistance on comput er matters.
I thank Ilya Mandel, my long-time roommate, for his forbeara nce regarding my poor
housekeeping abilities and tendency to consume his supplie s, as well as for many interesting
conversations and a memorable roadtrip from Pasadena to San Jos´ e, Costa Rica.
I thank Jie Yang, with whom I worked as a teaching assistant fo r two years, for her
superhuman efficiency, sunny disposition, and willingness t o take on more than her share
of the work.
I thank various Irishmen for arguments, and Anura Abeyesing he, Lotty Ackermann,
Christian Bauer, Xavier Calmet, Chris Lee, Sonny Mantry, Mi chael Salem, Graeme Smith,
Ben Toner, Lisa Tracy, and other members of my class and my res earch group whom I was
privileged to know personally.
I thank Jacob Bourjaily, Oleg Evnin, Jernej Kamenik, David M aybury, Brian Murray,
Jon Pritchard, Ketan Vyas, and other students with whom I had occassion to discuss
physics.
I thank physicists Nima Arkani-Hamed, J. D. Bjorken, Roman B uniy, Andy Frey, Jaume
Garriga, Holger Gies, Walter Goldberger, Jim Isenberg, Ted Jacobson, Marc Kamionkowski,
Alan Kosteleck´ y, Anton Kapustin, Eric Linder, Juan Maldac ena, Eugene Lim, Ian Low,
Guy Moore, Lubos Motl, Yoichiru Nambu, Hiroshi Ooguri, Kris hna Rajagopal, Michael
Ramsey-Musolf, John Schwarz, Matthew Schwartz, Guy de T´ er amond, Kip Thorne, and
Alex Vilenkin, for questions, comments, and discussions.
I thank Richard Berg, David Berman, Ed Creutz, John Dlugosz, Lars Falk, Monwhea
Jeng, Lewis Mammel, Carl Mungan, Frederick Ross, Wolf Rueck ner, Tom Snyder, and the
other professional and amateur physicists who commented on the work in Chapter 6.
I thank the professors with whom I worked as a teaching assist ant, David Goodstein,
Marc Kamionkowski, Bob McKeown, and Mark Wise, for their pat ience and understanding.
vii
I thank my father, mother, and brother for their support and a dvice.
I thank Caltech for sustaining me as a Robert A. Millikan grad uate fellow (2001–2004)
and teaching assistant (2004–2006). I was also supported du ring the summer of 2005 as a
graduate research associate under the Department of Energy contract DE-FG03-92ER40701.
viii
Abstract
We begin by reviewing our current understanding of massless particles with spin 1 and spin
2 as mediators of long-range forces in relativistic quantum field theory. We discuss how a
description of such particles that is compatible with Loren tz covariance naturally leads to
a redundancy in the mathematical description of the physics , which in the spin-1 case is
local gauge invariance and in the spin-2 case is the diffeomor phism invariance of General
Relativity. We then discuss the Weinberg-Witten theorem, w hich further underlines the
need for local invariance in relativistic theories with mas sless interacting particles that have
spin greater than 1/2.
This discussion leads us to consider a possible class of mode ls in which long-range in-
teractions are mediated by the Goldstone bosons of spontane ous Lorentz violation. Since
the Lorentz symmetry is realized non-linearly in the Goldst ones, these models evade the
Weinberg-Witten theorem and could potentially also evade t he need for local gauge invari-
ance in our description of fundamental physics. In the case o f gravity, the broken symmetry
would protect the theory from having non-zero cosmological constant, while the composite-
ness of the graviton could provide a solution to the perturba tive non-renormalizability of
linear gravity.
This leads us to consider the phenomenology of spontaneous L orentz violation and
the experimental limits thereon. We find the general low-ene rgy effective action of the
Goldstones of this kind of symmetry breaking minimally coup led to the usual Einstein
gravity and we consider observational limits resulting fro m modifications to Newton’s law
and from gravitational ˇCerenkov radiation of the highest-energy cosmic rays. We co mpare
this effective theory with the “ghost condensate” mechanism , which has been proposed in
the literature as a model for gravity in a Higgs phase.
Next, we summarize the cosmological constant problem and co nsider some issues related
to it. We show that models in which a scalar field causes the sup er-acceleration of the
ix
universe generally exhibit instabilities that can be more b roadly connected to the violation
of the null-energy condition. We also discuss how the equati on of state parameter w=p/ρ
evolves in a universe where the dark energy is caused by a ghos t condensate. Furthermore,
we comment on the anthropic argument for a small cosmologica l constant and how it is
weakened by considering the possibility that the size of the primordial density perturbations
created by inflation also varies over the landscape of possib le universes.
Finally, we discuss a problem in elementary fluid mechanics t hat had eluded a definitive
treatment for several decades: the reverse sprinkler, comm only associated with Feynman.
We provide an elementary theoretical description compatib le with its observed behavior.
x
Contents
Acknowledgments v
Abstract viii
1 Introduction 1
1.1 Notation and conventions . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2
2 Massless mediators 5
2.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5
2.1.1 Unbearable lightness . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
2.1.2 Overview . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8
2.2 Polarizations and the Lorentz group . . . . . . . . . . . . . . . . . . . . . . 8
2.2.1 The little group . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9
2.2.2 Massive particles . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1 0
2.2.3 Massless particles . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12
2.3 The vector field . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1 3
2.3.1 Vector field with j= 0 . . . . . . . . . . . . . . . . . . . . . . . . . . 15
2.3.2 Vector field with j= 1 . . . . . . . . . . . . . . . . . . . . . . . . . . 15
2.3.3 Massless j= 1 particles . . . . . . . . . . . . . . . . . . . . . . . . . 17
2.4 Why local gauge invariance? . . . . . . . . . . . . . . . . . . . . . . . . . . . 19
2.4.1 Expecting the Higgs . . . . . . . . . . . . . . . . . . . . . . . . . . . 20
2.4.2 Further successes of gauge theories . . . . . . . . . . . . . . . . . . . 21
2.5 Massless j= 2 particles and diffeomorphism invariance . . . . . . . . . . . . 22
2.6 The Weinberg-Witten theorem . . . . . . . . . . . . . . . . . . . . . . . . . 23
2.6.1 The j >1/2 case . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23
2.6.2 The j >1 case . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25
xi
2.6.3 Why are gluons and gravitons allowed? . . . . . . . . . . . . . . . . 25
2.6.4 Gravitons in string theory . . . . . . . . . . . . . . . . . . . . . . . . 28
2.7 Emergent gravity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29
3 Goldstone photons and gravitons 32
3.1 Emergent mediators . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 32
3.2 Nambu and Jona-Lasinio model (review) . . . . . . . . . . . . . . . . . . . . 37
3.3 An NJL-style argument for breaking LI . . . . . . . . . . . . . . . . . . . . 41
3.4 Consequences for emergent photons . . . . . . . . . . . . . . . . . . . . . . . 50
4 Phenomenology of spontaneous Lorentz violation 52
4.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 52
4.2 Phenomenology of Lorentz violation by a background sour ce . . . . . . . . . 54
4.3 Effective action for the Goldstone bosons of spontaneous Lorentz violation . 56
4.4 The long-range gravitational preferred-frame effect . . . . . . . . . . . . . . 58
4.5 A cosmic solid . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 67
5 Some considerations on the cosmological constant problem 70
5.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 70
5.2 Gradient instability for scalar models of the dark energ y withw<−1 . . . 73
5.3 Time evolution of wfor ghost models of the dark energy . . . . . . . . . . . 77
5.4 Anthropic distribution for Λ and primordial density per turbations . . . . . 78
6 The reverse sprinkler 88
6.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 88
6.2 Pressure difference and momentum transfer . . . . . . . . . . . . . . . . . . 90
6.3 Conservation of angular momentum . . . . . . . . . . . . . . . . . . . . . . 93
6.4 History of the reverse sprinkler problem . . . . . . . . . . . . . . . . . . . . 97
6.5 Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1 01
Bibliography 103
xii
List of Figures
2.1 Feynman diagram for scattering mediated by scalar field . . . . . . . . . . . . 7
2.2 Schematic of Witten’s argument against an emergent theo ry of gravity . . . . 30
3.1 Diagrammatic Schwinger-Dyson equation . . . . . . . . . . . . . . . . . . . . 38
3.2 Primed self-energy in a theory with a four-fermion inter action . . . . . . . . 39
3.3 Fermion and antifermion energies at finite densities . . . . . . . . . . . . . . 42
3.4 Four-fermion vertex as two massive, zero momentum photo n exchanges . . . 44
3.5 Plots self-consistent equations for the fermion mass m. . . . . . . . . . . . . 46
3.6 Further plots of self-consistent equations for the ferm ion massm. . . . . . . 49
3.7 Radiative corrections for the effective potential of the auxiliary field Aµ. . . 50
3.8 Graphic representation of how radiative corrections gi ve a finite ∝an}b∇acketle{tAµ∝an}b∇acket∇i}ht. . . . 51
4.1 Test mass orbiting a source moving with respect to the pre ferred frame . . . 66
4.2 Modification to gravity by perturbations in the CMB . . . . . . . . . . . . . 67
5.1 Tadpole diagram corresponding to the cosmological cons tant term . . . . . . 71
5.2 Piston filled with vacuum energy . . . . . . . . . . . . . . . . . . . . . . . . . 72
5.3 Effective coupling of two gravitons to several quanta of t he scalar ghost field 74
6.1 Closed sprinkler in a tank . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 90
6.2 Open sprinkler in a tank . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 91
6.3 Fluid flow in a pressure gradient . . . . . . . . . . . . . . . . . . . . . . . . . 92
6.4 Force creating the flow into the reverse sprinkler . . . . . . . . . . . . . . . . 93
6.5 Tank recoiling as water rushes out of it . . . . . . . . . . . . . . . . . . . . . 94
6.6 Machine gun in a floating ship . . . . . . . . . . . . . . . . . . . . . . . . . . 95
6.7 Water flowing out of a shower head . . . . . . . . . . . . . . . . . . . . . . . 96
6.8 Illustrations from Ernst Mach’s Mechanik . . . . . . . . . . . . . . . . . . . . 98
1
Chapter 1
Introduction
est aliquid, quocumque loco, quocumque recessu
unius sese dominum fecisse lacertae.
— Juvenal, Satire III
He thought he saw a Argument
That proved he was the Pope:
He looked again, and found it was
A Bar of Mottled Soap.
“A fact so dread,” he faintly said,
“Extinguishes all hope!”
— Lewis Carroll, Sylvie and Bruno Concluded
This dissertation is essentially a collection of the variou s theoretical investigations that
I pursued as a graduate student and that progressed to a publi shable state. It is difficult,
a posteriori , to come up with a theme that will unify them all. Even the absu rdly broad
title that I have given to this document fails to account at al l for Chapter 6, which concerns
a long-standing problem in elementary fluid mechanics. Ther efore I will not attempt any
such artificial unification here.
I have made an effort, however, to make this thesis more than co llation of previously
published papers. To that end, I have added material that rev iews and clarifies the relevant
physics for the reader. Also, as far as possible, I have compl emented the previously published
research with discussions of recent advances in the literat ure and in my own understanding.
Chapter 2 in particular was written from scratch and is inten ded as a review of the
relationship between massless particles, Lorentz invaria nce (LI), and local gauge invariance.
In writing it I attempted to answer the charge half-seriousl y given to me as a first-year
graduate student by Mark Wise of figuring out why we religious ly follow the commandment
of promoting the global gauge invariance of the Dirac Lagran gian to a local invariance in
2
order to obtain an interacting theory. Consideration of the role of local gauge invariance in
quantum field theories (QFT’s) with massless, interacting p articles also helps to motivate
the research described in Chapter 3.
Chapter 3 brings up spontaneous Lorentz violation, which is the idea that perhaps the
quantum vacuum of the universe is not a Lorentz singlet (or, t o put it otherwise, that empty
space is not empty). The idea that gravity might be mediated b y the Goldstone bosons
of such a symmetry breaking is attractive because it offers a p ossible solution to two of
the greatest obstructions to a quantum description of gravi ty: the non-renormalizability of
linear gravity, and the cosmological constant problem.
The work described in Chapter 4 seeks to place experimental l imits on how large spon-
taneous Lorentz violation can be when coupled to ordinary gr avity. This line of research is
independent from the ideas of Chapter 3 and applies to a wide v ariety of models in which
cosmological physics takes place in a background that is not a Lorentz singlet.
Chapter 5 begins with a brief overview of the cosmological co nstant problem, one of
the greatest puzzles in modern theoretical physics. The nex t three sections of that chapter
concern original results that are connected to that problem . Section 5.2 in particular has
applications beyond the cosmological constant problem, as it offers a theorem that helps
connect the energy conditions of General Relativity (GR) wi th considerations of stability.
All of this work concerns both QFT and GR, our two most powerfu l (though mutually
incompatible) tools for describing the universe at a fundam ental level. In Chapter 6 we
consider an amusing problem about introductory college phy sics that, surprisingly, had
evaded a completely satisfactory treatment for several dec ades.
1.1 Notation and conventions
We work throughout in units in which /planckover2pi1=c= 1. Electrodynamical quantities are given in
the Heaviside-Lorentz system of units in which the Coloumb p otential of a point charge q
is
Φ =q
4πr.
We also work in the convention in which the Fourier transform and inverse Fourier
3
transform in ndimensions are
f(x) =/integraldisplaydnk
(2π)n/2˜f(k)e−ik·x;˜f(x) =/integraldisplaydnk
(2π)n/2f(x)eik·x.
Lorentz 4-vectors are written as x= (x0,x1,x2,x3), wherex0is the time component
andx1,x2, andx3are the ˆx,ˆy, and ˆzspace components respectively. Spatial vectors are
denoted by boldface, so that we also write x= (x0,x). Unit spatial vectors are denoted by
superscript hats. Greek indices such as µ,ν,ρ , etc. are understood to run from 0 to 3, while
Roman indices such as i,j,k, etc. are understood to run from 1 to 3. Repeated indices are
always summed over, unless otherwise specified.
We takegµνto represent the full metric in GR, while ηµν= diag(+1,−1,−1,−1) is the
Minkowski metric of flat space-time. Indices are raised and l owered with the appropriate
metric. The square of a tensor denotes the product of the tens or with itself, with all the
indices contracted pairwise with the metric. Thus, for inst ance, the d’Alembertian operator
in flat spacetime is
/square=−∂2=−∂µ∂µ=−ηµν∂µ∂ν=−∂2
0+∇2.
We define the Planck mass as MPl=/radicalbig
1/8πG, whereGis Newton’s constant. For linear
gravity we expand the metric in the form gµν=ηµν+M−1
Plhµνand keep only terms linear
inh. In Chapter 2 we will work in units in which MPl= 1. Elsewhere we will show the
factors ofMPlexplicitly.
We use the chiral basis for the Dirac matrices
γµ=
0σµ
¯σµ0
, γ5=
−1 0
0 1
,
whereσµ= (1,σ), ¯σµ= (1,−σ), and theσi’s are the Pauli matrices
σ1=
0 1
1 0
, σ2=
0−i
i0
, σ3=
1 0
0−1
.
All other conventions are the standard ones in the literatur e.
In writing this thesis, I have used the first person plural (“w e”) whenever discussing
4
scientific arguments, regardless of their authorship. I hav e used the first person singular
only when referring concretely to myself in introductory of parenthetical material. I feel
that this inconsistency is justified by the avoidance of styl istic absurdities.
5
Chapter 2
Massless mediators
Did he suspire, that light and weightless down
perforce must move.
— William Shakespeare, Henry IV, part ii , Act 4, Scene 3
You lay down metaphysic propositions which infer universal consequences,
and then you attempt to limit logic by despotism.
— Edmund Burke, Reflections on the Revolution in France
2.1 Introduction
I have sometimes been asked by scientifically literate layme n (my father, for instance, who is
a civil engineer, and my ophthalmologist) to explain to them how a particle like the photon
can be said to have no mass. How would a particle with zero mass be distinguishable from
no particle at all? My answer to that question has been that in modern physics a particle is
not defined as a small lump of stuff (which is the mental image im mediately conveyed by the
word, as well as the non-technical version of the classical d efinition of the term) but rather
as an excitation of a field, somewhat akin to a wave in an ocean. In that sense, masslessness
means something technical: that the excitation’s energy go es to zero when its wavelength
is very long. I have then added that masslessness also means t hat those excitations must
always propagate at the speed of light and can never appear to any observer to be at rest.
Here I will attempt a fuller treatment of this problem. Much o f the professional life of
a theoretical physicist consists of ignoring technical diffi culties and underlying conceptual
confusion, in the hope that something publishable and perha ps even useful might emerge
from his labor. If the theorist had to proceed in strictly log ical order, the field would advance
6
very slowly. But, on the other hand, the only thing that can ul timately protect us from
being seriously wrong is sufficient clarity about the basics. In modern physics, long-range
forces (electromagnetism and gravity) are understood to be mediated by massless particles
with spinj≥1. The description of such massless particles in quantum fiel d theory (QFT)
is therefore absolutely central to our current understandi ng of nature.
Therefore, I have decided to use the opportunity afforded by t he writing of this thesis to
review the subject. My goals are to elucidate why a relativis tic description of massless parti-
cles with spin j≥1 naturally requires something like local gauge invariance (which is not a
physical symmetry at all, but a mathematical redundancy in t he description of the physics)
and to clarify under what circumstances one might expect to e vade this requirement.
I shall conclude with a discussion of how these consideratio ns apply to whether some
of the major outstanding problems of quantum gravity could b e addressed by considering
gravity to be an emergent phenomenon in some theory without f undamental gravitons.
Nothing in this chapter will be original in the least, but it w ill provide a motivation for
some of the original work presented in Chapter 3.
2.1.1 Unbearable lightness
In his undergraduate textbook on particle physics, David Gr iffiths points out that massless
particles are meaningless in Newtonian mechanics because t hey carry no energy or momen-
tum, and cannot sustain any force. On the other hand, the rela tivistic expression for energy
and momentum:
pµ= (E,p) =γm(1,v) (2.1)
allows for non-zero energy-momentum for a massless particl e ifγ≡/parenleftbig
1−v2/parenrightbig−1/2→ ∞,
which requires |v| →1. Equation (2.1) doesn’t tell us what the energy-momentum i s, but
we assume that the relation p2=m2is valid form= 0, so that a massless particle’s energy
Eand momentum pare related by
E=|p|. (2.2)
Griffiths adds that
Personally I would regard this “argument” as a joke, were it n ot for the fact
that [massless particles] are known to exist in nature. They do indeed travel at
7/A1
Figure 2.1: Feynman diagram for the scattering of two particles that int eract through the exchange
of a mediator.
the speed of light and their energy and momentum arerelated by [Eq. (2.2)]
([1]).
The problem of what actually determines the energy of the mas sless particle is solved not
by special relativity, but by quantum mechanics, via Planck ’s formulaE=ω, whereωis an
angular frequency (which is an essentially wave-like prope rty). Thus massless particles are
the creatures of QFT par excellence , because, at least in current understanding, they can
only be defined as relativistic, quantum-mechanical entiti es. Like other subjects in QFT,
describing massless particles requires arguments that wou ld seem absurd were it not for the
fact that they yield surprisingly useful results that have g iven us a handle on observable
natural phenomena.
We need massless particles because we regard interaction fo rces as resulting from the
exchange of other particles, called “mediators.” Figure 2. 1 shows the Feynman diagram that
represents the leading perturbative term in the amplitude f or the scattering of two particles
(represented by the solid lines) that interact via the excha nge of a mediator (represented
by the dashed line). We can calculate this Feynman diagram in QFT and match the result
to what we would get in non-relativistic quantum mechanics f rom an interaction potential
V(r) (see, e.g., Section 4.7 in [2]). The result is
V(r) =−g2
4πe−µr
r, (2.3)
wheregis the coupling constant that measures the strength of the in teraction and µis the
mediator’s mass. Therefore, a long-range force requires µ= 0. In order to accommodate
the observed properties of the long-range electromagnetic and gravitational interactions, we
also need to give the mediator a on-zero spin. We will see that this is non-trivial.
8
2.1.2 Overview
In this chapter we shall first briefly review how one-particle states are defined in QFT
and how their polarizations correspond to basis states in ir reducible representations of the
Lorentz group. We will emphasize the difference between the c ase when the mass mof the
particle is positive and the case when it is zero. We shall pro ceed to use these tools to build a
fieldAµthat transforms as a Lorentz 4-vector, first for m>0 and then for m= 0. We shall
conclude that the relativistic description of a massless sp in-1 field requires the introduction
of local gauge invariance. Similarly, we will point out how t he relativistic description of a
massless spin-2 particle that transforms like a two-index L orentz tensor requires something
like diffeomorphism invariance (the fundamental symmetry o f GR). Our discussion of these
matters will rely heavily on the treatment given in [3].
We will then seek to formulate a solid understanding of the me aning of local gauge
invariance and diffeomorphism invariance as redundancies o f the mathematical description
required to formulate a relativistic QFT with massless medi ators. To this end we will also
review the Weinberg-Witten theorem ([4]) and conclude by co nsidering how it might be
possible to do without gauge invariance and evade the Weinbe rg-Witten theorem in an
attempt to write a QFT of gravity without UV divergences.
2.2 Polarizations and the Lorentz group
We define one-particle states to be eigenstates of the 4-mome ntum operator Pµand label
them by their eigenvalues, plus any other degrees of freedom that may characterize them:
Pµ|p,r∝an}b∇acket∇i}ht=pµ|p,r∝an}b∇acket∇i}ht. (2.4)
Under a Lorentz transformation Λ that takes pto Λp, the state transforms as
|p,r∝an}b∇acket∇i}ht →U(Λ)|p,r∝an}b∇acket∇i}ht (2.5)
whereU(Λ) is a unitary operator in some representation of the Loren tz group. The 4-
momentum itself transforms in the fundamental representat ion, so that
U†(Λ)PµU(Λ) = Λµ
νPν. (2.6)
9
The 4-momentum of the transformed state is therefore given b y
PµU(Λ)|p,r∝an}b∇acket∇i}ht=U(Λ)/bracketleftig
U†(Λ)PµU(Λ)/bracketrightig
|p,r∝an}b∇acket∇i}ht=U(Λ)Λµ
νpν|p,r∝an}b∇acket∇i}ht= (Λp)µU(Λ)|p,r∝an}b∇acket∇i}ht,(2.7)
which implies that U(Λ)|p,r∝an}b∇acket∇i}htmust be a linear combination of states with 4-momentum Λ p:
U(Λ)|p,r∝an}b∇acket∇i}ht=/summationdisplay
r′crr′(p,Λ)/vextendsingle/vextendsingleΛp,r′/angbracketrightbig
. (2.8)
If the matrix crr′(p,Λ) in Eq. (2.8), for some fixed p, is written in block-diagonal form, then
each block gives an irreducible representation of the Loren tz group. We will call particles
in the same irreducible representation “polarizations.” T he number of polarizations is the
dimension of the corresponding irreducible representatio n.1
2.2.1 The little group
For a particle with mass given by m=/radicalbig
p2≥0, let us choose an arbitrary reference 4-
momentum ksuch thatk2=m2. Any 4-momentum with the same invariant norm can be
written as
pµ=K(p)µ
νkν(2.9)
for some appropriate Lorentz transformation K(p).
Let us then define the “little group” as the group of Lorentz tr ansformations Ithat
leaves the reference kµinvariant:
Iµ
νkν=kµ. (2.10)
Then Eq. (2.8) can be approached by considering Drr′(I) =crr′(p=k,Λ =I) so that
U(I)|k,r∝an}b∇acket∇i}ht=/summationdisplay
r′Drr′(I)/vextendsingle/vextendsinglek,r′/angbracketrightbig
(2.11)
and defining 1-particle states with other 4-momenta by:
|p,r∝an}b∇acket∇i}ht=N(p)U(K(p))|k,r∝an}b∇acket∇i}ht, (2.12)
1Notice that in this choice of language a Dirac fermion has fou r polarizations: the spin-up and spin-down
fermion, plus the spin-up and spin-down antifermion.
10
whereN(p) is a normalization factor. If we impose that
∝an}b∇acketle{tk′,r′|k,r∝an}b∇acket∇i}ht=δr′rδ3(k′−k) (2.13)
for states with 4-momentum k, then
∝an}b∇acketle{tp′,r′|p,r∝an}b∇acket∇i}ht=N∗(p′)N(p)/angbracketleftbig
k′,r′/vextendsingle/vextendsingleU†(K(p′))U(K(p))/vextendsingle/vextendsinglek,r/angbracketrightbig
=N∗(p′)N(p)Dr′r/parenleftbig
K−1(p′)K(p)/parenrightbig
δ3(k′−k). (2.14)
Since theδ-function in the second line vanishes unless k′=k, this implies that the overlap
is zero unless p′=p, and theDmatrix in Eq. (2.14) is therefore trivial:
∝an}b∇acketle{tp′,r′|p,r∝an}b∇acket∇i}ht=|N(p)|2δr′rδ3(k′−k). (2.15)
We wish to rewrite Eq. (2.15) in terms of δ3(p′−p), to which we have argued it must be
proportional. It is not difficult to show that d3p/p0is a Lorentz-invariant measure when in-
tegrating on the mass shell p0=/radicalbig
p2+m2. This implies that δ3(k′−k) =δ3(p′−p)p0/k0
and we therefore have that
∝an}b∇acketle{tp′,r′|p,r∝an}b∇acket∇i}ht=|N(p)|2δr′rδ3(p′−p)p0/k0. (2.16)
Equation (2.16) naturally leads to the choice of normalizat ion
N(p) =/radicalbig
k0/p0. (2.17)
2.2.2 Massive particles
A massive particle will always have a rest frame in which its 4 -momentum is kµ= (m,0,0,0).
This is, therefore, the natural choice of reference 4-momen tum. It is easy to check that the
little group is then SO(3), which is the subgroup of the Lorentz group that includes only
rotations.
The generators of SO(3) may be written as
Ji=iǫijkxj∂k, (2.18)
11
which are the angular momentum operators and which obey the c ommutation relation
/bracketleftbig
Ji,Jj/bracketrightbig
=iǫijkJk. (2.19)
The Lie algebra of SO(3) is the same as that of SU(2), because both groups look identical in
the neighborhood of the identity. In quantum mechanics, the intrinsic angular momentum
of a particle (its spin) is a label of the dimensionality of th e representation of SU(2) that
we assign to it. A particle of spin jlives in the 2 j+1 dimensional representation of SU(2).
The generators of SO(1,3) may be written as
Jµν=i(xµ∂ν−xν∂µ), (2.20)
which are clearly anti-symmetric in the indices and which ob ey the commutation relation
[Jµν,Jρσ] =i(ηνρJµσ−ηµρJνσ−ηνσJµρ+ηµσJνρ). (2.21)
We may write the six independent components of Jµνas two three-component vectors:
Ki=J0i;Li=1
2ǫijkJjk, (2.22)
where Kis the generator of boosts and Lis the generator of rotations. Using Eqs. (2.21)
and (2.22), one can immediately show that these satisfy the c ommutation relations:
/bracketleftbig
Li,Lj/bracketrightbig
=iǫijkLk;/bracketleftbig
Li,Kj/bracketrightbig
=iǫijkKk;/bracketleftbig
Ki,Kj/bracketrightbig
=−iǫijkJk. (2.23)
Let us define two new 3-vectors:
J±=1
2(L±iK). (2.24)
Using Eq. (2.23) we can write their commutators as
/bracketleftig
Ji
±,Jj
±/bracketrightig
=iǫijkJk
±;/bracketleftig
Ji
±,Jj
∓/bracketrightig
= 0. (2.25)
That is, both J+andJ−separately satisfy the commutation relation for angular mo -
12
mentum, and they also commute with each other. This means tha t we can identify all
finite-dimensional representations of the Lorentz group SO(1,3) by pair of integer or half-
integer spins ( j+,j−) that correspond to two uncoupled representations of SO(3). The
Lorentz-transformation property of a left-handed Weyl fer mionψLcorresponds to (1 /2,0),
while (0,1/2) corresponds to the right-handed Weyl fermion ψR. A massive Dirac fermion
corresponds to the representation (1 /2,0)⊕(0,1/2).
A Lorentz 4-vector (that is, a quantity that transforms unde r the fundamental repre-
sentation of SO(3,1)), corresponds to (1 /2,1/2). This indicates that it can be decomposed
into a spin-1 and a spin-0 part, since 1 /2⊗1/2 = 1⊕0. Or, to put it otherwise, a general
Lorentz vector has four independent components, three of wh ich may be matched to the
three polarizations of a j= 1 particle and one to the single polarization of a j= 0 particle.
2.2.3 Massless particles
Since a massless particle has no rest frame, the simplest ref erence 4-momentum is k= (1,0,0,1).
The corresponding little group clearly contains as a subgro up rotations about the z-axis.
The little group can be parametrized as
I(δ,η,φ)µ
ν= Λ(δ,η)µ
ρΛ(φ)ρ
ν, (2.26)
where
Λ(φ)µ
ν=
1 0 0 0
0 cosφsinφ0
0−sinφcosφ0
0 0 0 1
(2.27)
and
Λ(δ,η)µ
ν=
1 +ζ δ η −ζ
δ1 0 −δ
η0 1 −η
ζ δ η 1−ζ
, (2.28)
withζ=/parenleftbig
δ2+η2/parenrightbig
/2.
13
It can be readily checked that
Λ(δ1,η1)µ
ρΛ(δ2,η2)ρ
ν= Λ(δ1+δ2,η1+η2)µ
ν, (2.29)
which implies that the little group is isomorphic to the grou p of rotations (by an angle φ)
and translations (by a vector ( δ,η)) in two dimensions.2UnlikeSO(3), this group, ISO(2),
is not semi-simple, i.e., it has invariant abelian subgroup s: the rotation subgroup defined
by Eq. (2.27) and the translation subgroup defined by Eq. (2.2 8).
This leads to the important consequence that massless one-p article states |p,r∝an}b∇acket∇i}htcan have
only two polarization, called “helicities,” given by the co mponent of the angular momentum
along its direction of motion. The physical reason for this i s that only the angular momen-
tum component associated with the rotations in Eq. (2.27) ca n define discrete polarizations.
Helicities are Lorentz-invariant, unlike the polarizatio ns of a massive particle.
It is clear that massless particles in QFT are different from m assive ones. It is possible to
understand some of the properties of massless particles by c onsidering them as massive and
then taking the m→0 limit carefully, but this discussion should make it appare nt that this
limiting procedure is fraught with danger. We shall explore this issue in the construction
of the vector field.
2.3 The vector field
We seek a causal, free quantum field Aµthat transforms like a Lorentz 4-vector. By analogy
to the procedure used to obtain free quantum fields with spin 0 and 1/2 (see, e.g., Chapters
2 and 3 in [2], or Sections 5.2 to 5.5 in [3]), we start by writin g
Aµ(x) =/integraldisplayd3p
(2π)3/2/summationdisplay
r/parenleftig
ǫµ
r(p)ar(p)eip·x+ǫµ∗
r(p)a†
r(p)e−ip·x/parenrightig
, (2.30)
where the index rruns over the physical polarizations of the field, while aanda†are
the creation and destruction operators for particles of the corresponding momentum and
polarization that obey bosonic commutation relations, and pµ=/parenleftig/radicalbig
m2+p2,p/parenrightig
.
LetK(p) be the Lorentz transformation (boost) that takes a particl e of massmfrom
2The Lorentz transformation in Eq. (2.28) is, of course, not a physical translation. It just happens that
the group of such matrices is isomorphic to the group of trans lations on the plane.
14
rest to a 4-momentum p. It can be shown that the measure d3p/p0is Lorentz-invariant
when integrating on the mass-shell p2=m2. Since both pµandAµare Lorentz 4-vectors,
we must have
ǫµ
r(p) =/radicalbiggm
p0K(p)µ
νǫν
r(0). (2.31)
Now consider the behavior of ǫµ
r(0) under an infinitesimal rotation. For our field Aµ(x)
in Eq. (2.30) to have a definite spin j, we must have that
Lµ
νǫν
r(0) = S(j)
rr′ǫµ
r′(0), (2.32)
where the three components of S(j)are the standard spin matrices for spin j. Equation
(2.32) follows immediately from requiring ǫµ
r(0) to transform under rotations as both a
4-vector and as a spin- jobject.3
For the rotation generators in the fundamental representat ion ofSO(1,3) we have:
(Li)0
0= (Li)0
j= (Li)j
0= 0, (2.33)
(Li)jk=iǫi
jk. (2.34)
Therefore, for ( L2)µ
ν=/summationtext
i(Li)µ
ρ(Li)ρ
ν, we have
(L2)0
0= (L2)j
0= (L2)0
j= 0 ; ( L2)j
k= 2ηj
k. (2.35)
Meanwhile, recall that, for the spin matrices,
(S(j) 2)rr′=j(j+ 1)δrr′. (2.36)
Using Eqs. (2.32), (2.35), and (2.36) we therefore obtain th at
ǫi
r(0) =j(j+ 1)
2ǫi
r(0) ;j(j+ 1)ǫ0
r(0) = 0. (2.37)
Equation (2.37), combined with Eq. (2.31), leaves us only tw o posibilities if the field
Aµ(x) in Eq. (2.30) is to transform as a 4-vector:
3It should perhaps also be pointed out that in Eq. (2.32) the in dices µ, νin the left-hand side indicate
components of the three matrices Lidefined in Eq. (2.22). In Eq. (2.21) µ, νlabeled the matrices themselves.
15
•Eitherj= 0 andǫ0(0) is the only non-vanishing component,
•orj= 1 and the three ǫi(0)’s are the only non-vanishing components
This agrees with the claim made at the end of the previous sect ion, which we had based on
1/2⊗1/2 = 1⊕0. Let us explore both possibilities.
2.3.1 Vector field with j= 0
For thej= 0 case we can chose the conventionally normalized ǫ0(0) =i/radicalbig
m/2, which, by
Eq. (2.31) gives
ǫµ(p) =ipµ/radicalbigg
1
2p0. (2.38)
One can then compare the resulting form for Aµ(x) in Eq. (2.30) to the form for a free
scalar field and conclude that this vector field has the form
Aµ(x) =∂µφ(x) (2.39)
forφ(x) a free, Lorentz scalar field. Notice that as the field φhas a single physical polar-
ization, so also does Aµ, and that even though our construction of the vector field ass umed
anm>0 in Eq. (2.31), the m→0 limit in this case is perfectly sensible.4
2.3.2 Vector field with j= 1
Now consider the case where the vector field has j= 1. Following the popular convention
we write
ǫµ
r=±1(0) =∓1
2√m(ηµ
1±iηµ
2) (2.40)
and
ǫµ
r=0(0) =/radicalbigg
1
2mηµ
3. (2.41)
We may check that the raising and lowering operators S(1)
±=S(1)
1±iS(1)
2act appropriately
on these polarization vectors. For a plane-wave propagatin g along the i= 3 spatial direction,
r=±1 correspond to two transverse, circular polarizations of t he vector field, while r= 0
corresponds to the longitudinal polarization.
4This kind of massless, spinless vector field will appear agai n in the discussion of the “ghost condensate”
mechanism in Chapter 4.
16
We may rewrite the field Aµin terms of polarization vectors that are mass-independent
by introducing
˜ǫµ
r(0) =√
2mǫµ
r(0) (2.42)
then we have that Eq. (2.30) becomes
Aµ(x) =/integraldisplayd3p
(2π)3/21/radicalbig
2p01/summationdisplay
r=−1/parenleftig
˜ǫµ
r(p)ar(p)eip·x+ ˜ǫµ
r(p)a†
r(p)e−ip·x/parenrightig
, (2.43)
where ˜ǫµ
r(p) =K(p)µ
ν˜ǫν
r(0). The field in Eq. (2.43) obeys the equation of motion
/parenleftbig/square−m2/parenrightbig
Aµ(x) = 0. (2.44)
Notice also that
pµ˜ǫµ
r(p) =pµKµ
ν(p)˜ǫν
r(0) =/parenleftbig
K−1(p)p/parenrightbig
ν˜ǫν
r(0) =m˜ǫ0
r(0) = 0 (2.45)
implies that
∂µAµ= 0. (2.46)
In the limit m→0 the boost K(p) becomes the identity and ˜ ǫµ
r(p) = ˜ǫµ
r(0) for all p.
The field then obeys both /squareAµ= 0 and∂µAµ= 0.5The fact that there are complications
in this limit is revealed by using Eq. (2.31) and the form of ˜ ǫµ
r(0)’s to obtain
Πµν(p)≡1/summationdisplay
r=−1˜ǫµ
r(p)˜ǫν
r(p) =ηµν+pµpν
m2. (2.47)
Notice that Πµνpν= 0, while Πµνkν=kµfork·p= 0, which means Πµνis a projection
unto the space orthogonal to pµ. Equation (2.47) clearly is not finite as m→0. This will
be a problem if we try to directly couple Aµto anything in a Lorentz-invariant way,
Lint∝Aµjµ, (2.48)
because then the rate at which Aµ’s would be emitted by the interaction would be propor-
5Therefore taking the m→0 limit of the spin-1 vector field automatically gives us the m assless field in
the Lorenz gauge.
17
tional to
/summationdisplay
r|˜ǫµ
r(p)∝an}b∇acketle{tjµ∝an}b∇acket∇i}ht|2= Πµν(p)∝an}b∇acketle{tjµ∝an}b∇acket∇i}ht∝an}b∇acketle{tjν∝an}b∇acket∇i}ht∗=∝an}b∇acketle{tjµ∝an}b∇acket∇i}ht∝an}b∇acketle{tjµ∝an}b∇acket∇i}ht∗+1
m2|p· ∝an}b∇acketle{tj∝an}b∇acket∇i}ht|2, (2.49)
which clearly diverges as m→0 unless we impose that p· ∝an}b∇acketle{tj∝an}b∇acket∇i}ht= 0. That is, in the presence
of an interaction of the form Eq. (2.48), we must require that the current to which the field
couples be conserved,
∂µ∝an}b∇acketle{tjµ∝an}b∇acket∇i}ht= 0, (2.50)
in order to avoid an infinite rate of emission.
As emphasized earlier in this chapter, the spin of a massless particle must point either
parallel or anti-parallel to its direction of propagation. These possibilities correspond to the
longitudinal polarizations ˜ ǫµ
±1. A massless particle cannot have a longitudinal polarizati on
˜ǫµ
0. The requirement of current conservation in Eq. (2.50) ensu res that the longitudinal
polarization decouples from the current jµin them→0 limit, so that it cannot be produced
by the interaction in Eq. (2.48).
2.3.3 Massless j= 1particles
Let us now try to construct a genuinely massless vector field w ith non-zero spin j. To
that effect we adopt an arbitrary reference momentum k= (0,0,1) and a corresponding
light-like reference 4-momentum k= (1,0,0,1). LetK(p) be now defined as the Lorentz
transformation that takes a massless particle with referen ce momentum kto a general
momentum p. We can write this transformation as the composition of a rot ation (from the
direction of kto the direction of p) followed by a boost along the direction of pthat scales
the magnitude. Then
ǫµ
r(p) =K(p)µ
νǫν
r(k). (2.51)
We now require that ǫµ
r(k) transform as both a massless particle with helicity r=±j
and as a 4-vector. For rotations by an angle φaround the axis of k, we must have
eirφǫµ
r(k) = Λ(φ)µ
νǫν
r(k), (2.52)
where Λ(φ)µ
νis the Lorentz transformation matrix corresponding to the r otation, given in
18
Eq. (2.27). For Eq. (2.52) to be true of a general φin thej= 1 case, we must have
ǫµ
±1(k)∝(0,1,±i,0) (2.53)
and we might as well normalize this solution to match the ˜ ǫµ
r’s in Eq. (2.42), giving
ǫµ
±1(k) =1√
2(0,1,±i,0). (2.54)
These are the same polarization vectors that we obtained pre viously in the m→0 limit of
the massive vector field.
But the little group for massless particles is larger than th eO(2) =U(1) group rep-
resented by Eq. (2.27), as was seen in Subsection 2.2.3. For o ur field to transform as a
4-vector we would also require that
ǫµ
r(k) = Λ(δ,η)µ
νǫν
r(k), (2.55)
where Λ(δ,η)µ
νwas given in Eq. (2.28). Plugging in the polarization 4-vect ors in Eq. (2.54)
we can see immediately that this is impossible because, unde r the transformation Λ( δ,η),
ǫµ
±1(k)→ǫµ
±1(k) +kµ
|k|δ±iη√
2. (2.56)
Thus we are forced to accept that the one-particle states of a massless spin-1 vector field
are not Lorentz-covariant under the action of their little g roup, but only covariant up to
a term proportional to the reference kµ. If we then construct the general states using Eq.
(2.51) and
Aµ(x) =/integraldisplayd3p
(2π)3/21/radicalbig
2p0/summationdisplay
r=±1/parenleftig
ǫµ
r(p)ar(p)eip·x+ǫµ
r(p)a†
r(p)e−ip·x/parenrightig
(2.57)
we see that we are forced to accept that Aµ(x) transforms under a general Lorentz trans-
formation Λ as:
Aµ(x)→Λµ
νAν(Λx) +∂µΩ(x,Λ) (2.58)
where Ω is some function of the coordinates xand the parameters of the Lorentz transfor-
mation Λ.
19
Equation (2.58) should, in my opinion, be regarded as a disas ter. Massless spin-1 quan-
tum fields, which we need in order to explain the observed prop erties of the electromagnetic
interaction, are incompatible with one of the most sacred pr inciples of modern physics:
Lorentz covariance.6It is not, however, an irretrievable disaster, and in fact th ere will be a
rich silver lining to it.
We can “save” Lorentz covariance by announcing that two field s related by the trans-
formation
Aµ→Aµ+∂µΩ (2.59)
describe the same physics, so that the second term in Eq. (2.5 8) becomes irrelevant.7We
can couple such an Aµif the interaction is of the form Lint∝Aµjµfor a conserved current
jµ, because in that case the coupling is invariant under transf ormations of the form in Eq.
(2.59). Notice that this requirement on the coupling of Aµagrees with what we imposed
earlier, by Eq. (2.49), in order to avoid an infinite rate of em ission for the vector field in
them→0 limit.
It is easy to construct a genuinely Lorentz-covariant two-i ndex field strength tensor that
is invariant under Eq. (2.59):
Fµν=∂µAν−∂νAµ. (2.60)
Lorentz-invariant couplings to this field strength would be gauge-invariant, but the presence
of derivatives in Eq. (2.60) means that the resulting forces must fall off faster with distance
than an inverse-square law (i.e., they cannot be long-range forces).
2.4 Why local gauge invariance?
The Dirac Lagrangian for a free fermion, L=¯ψ(i∂ /−m)ψis invariant under the global
U(1) gauge transformation ψ→eiαψ. This global symmetry, by Noether’s theorem, implies
6This statement may seem peculiar in light of the fact that the Lorentz group was first discovered as the
symmetry of the Maxwell equations of classical electrodyna mics. But those equations are written in terms of
the fields EandB. The scalar and vector potentials ( A0andArespectively) enter classical electrodynamics
only as computational aids. It is quantum mechanics which requires a formulation in terms of Aµ.
7This irresistibly brings to my mind a scene from the Woody All en movie comedy Bananas in which
victorious rebel commander Esposito announces from the Pre sidential Palace that “from this day on, the
official language of San Marcos will be Swedish... Furthermor e, all children under 16 years old are now 16
years old.”
20
conservation of the current:
jµ=¯ψγµψ . (2.61)
In the established model of quantum electrodynamics, this L agrangian is transformed into
an interacting theory by making the gauge invariance local: The phaseαis allowed to be a
function of the space-time point x. This requires the introduction of a gauge field Aµwith
the the transformation property
Aµ→Aµ+∂µα (2.62)
and the use of a covariant derivative Dµ=∂µ−iAµinstead of the usual derivative ∂µ.
This procedure automatically couples Aµto the conserved current in Eq. (2.61) so that the
coupling is invariant under transformations of the form Eq ( 2.62). We than add a Lorentz-
invariant kinetic term −F2
µν/4 for the field Aµ. The generalization to non-abelian gauge
groups is well known, as is the Higgs mechanism to break the ga uge invariance spontaneously
and give the field Aµa mass.
This is what we are taught in elementary courses on QFT, but th e question remains:
Why do we promote a global symmetry of the free fermion Lagran gian to a local symmetry?
Equation (2.58) provides a deeper insight into the physical meaning of local gauge invariance:
a massless particle, having no rest frame, cannot have its sp in point along any axis other
than that of its motion. Therefore, it can have only two polar izations. By describing it as
a 4-vector, spin-1 field Aµ(which has three polarizations) a mathematical redundancy is
introduced.
This redundancy is local gauge invariance. A field with local gauge symmetry is coupled
to the conserved current of the corresponding global gauge s ymmetry in order to make the
coupling locally gauge-invariant. The procedure describe d of promoting the global gauge
symmetry to a local gauge invariance is therefore required i n order to couple fermions in a
Lorentz-invariant way via a long-range, spin-1 force.
2.4.1 Expecting the Higgs
Remarkably, local gauge invariance also comes to our aid in w riting sensible QFT’s for the
short-range weak nuclear interaction. At low energies, thi s interaction is naturally described
21
as being mediated by massive, spin-1 vector fields. The Lagra ngian for such a mediator must
look like
L=−1
4F2
µν+1
2m2A2−AµJµ, (2.63)
whereJµis the current to which it couples. But in the case of the weak n uclear interaction
this current is not conserved. At energy scales much higher t han themin Eq. (2.63), we
therefore expect the same problem we found in Subsection 2.3 .2 of a divergent emission rate
for the longitudinal polarization, unless other higher-de rivative operators, which were not
relevant at low energies, have come to our rescue.
In the standard model of particle physics, the resolution of this problem is to make
the mediators of the weak nuclear interaction gauge bosons, and then to break that gauge
invariance spontaneously by introducing a scalar Higgs fiel d with a non-zero VEV, thus
giving the bosons the mass that accounts for the short range o f the force they mediate. At
high energies the gauge invariance is restored. The problem atic longitudinal polarization
disappears and is transmuted into the Goldstone boson of the spontaneously broken sym-
metry. Since the Goldstone boson has no spin, it does not have the problem of a divergent
rate of emission. This is the reason why many billions of doll ars have been spent in the
search for that yet-unseen Higgs boson, a search soon to come to a head with the turning
on of the Large Hadron Collider (LHC) at CERN next year.
2.4.2 Further successes of gauge theories
Gauge theories as descriptions of the fundamental particle interactions have other very
attractive attributes. It was shown by ’t Hooft that these th eories are always renormalizable,
i.e., that the infinities that plague QFT’s can all be absorbe d into a redefinition of the bare
parameters of the theory, namely the masses and the coupling constants ([5]). Politzer
([6]) and, independently, Gross and Wilczek ([7]), showed t hat the renormalization flow
of the coupling constants in non-abelian gauge theories pro vides a natural explanation of
the observed phenomenon of asymptotic freedom, whereby the nuclear interactions become
more feeble at higher energies.
It is also widely believed, though not strictly demonstrate d, that QCD, the theory
in which the strong nuclear force is mediated by the bosons of anSU(3) gauge theory,
accounts for confinement, i.e., for the fact that the strongl y interacting fermions (quarks)
22
never occur alone and can appear only in bound states that are singlets ofSU(3). These
successes illustrate what we meant when we said in Subsectio n 2.3.3 that having to accept
local gauge symmetry was a disaster with a rich silver lining . For interesting accounts of
the history of local gauge invariance in classical and quant um physics, see [8, 9].
2.5 Massless j= 2particles and diffeomorphism invariance
We could repeat the sort of procedure used in Subsection 2.3. 3 in order to try to construct
Lorentz-covariant hµνout of the two helicities of a j= 2 massless field. This procedure
would similarly fail, requiring us to accept the transforma tion rule:
hµν(x)→Λµ
ρΛν
σhρσ(Λx) +∂µξν(x,Λ) +∂νξµ(x,Λ). (2.64)
Saving Lorentz covariance would then require announcing th at states related by a transfor-
mation of the form
hµν→hµν+∂µξν+∂νξµ(2.65)
are physically equivalent. We can construct a four-index fie ld strength tensor Rµνρσinvari-
ant under Eq. (2.65) that is anti-symmetric in µ,ν, anti-symmetric in ρ,σ, and symmetric
under exchange of the two pairs. But to accommodate a long-ra nge force we would need to
couplehµνto a quantity Θµνsuch that
∂µ∝an}b∇acketle{tΘµν∝an}b∇acket∇i}ht= 0. (2.66)
This Θµνis the stress-energy tensor obtained from translational in variance
xµ→xµ−ξµ, (2.67)
through Noether’s theorem.8Invariance under Eq. (2.65) corresponds to promoting the
translational symmetry in Eq. (2.67) to a local invariance b y lettingξµbe a function of x.
It turns out that the theory constructed in this way matches l inearized GR around a flat
background with hµνbeing the graviton field.
8If there were another conserved Θ′µν, there would have to be another conserved 4-vector besides pµ,
namely p′µ=/integraltext
d3xΘ′0µ. Kinematics would then allow only forward collisions.
23
It is well known that one can reconstruct the full GR uniquely from linear gravity by
a self-consistency procedure ([10, 11, 12]). Therefore a re lativistic QFT in flat spacetime
with a massless spin-2 particle mediating a long-range forc e essentially implies GR. In full
GR the invariance under Eq. (2.65) is a consequence of the inv ariance of the theory under
diffeomorphisms:
xµ→x′µ(x). (2.68)
Remarkably, we may therefore think of diffeomorphism invari ance as a redundancy required
by the relativistic description of a massless spin-2 partic le.
2.6 The Weinberg-Witten theorem
The Weinberg-Witten theorem9rules out the existence of massless particles with higher sp in
in a very wide class of QFT’s ([4]). In their original paper, t he authors present their elegant
proof very succinctly. This review is longer than the paper i tself, which may be justified
by the importance of this result in further clarifying the ne ed for local gauge invariance in
relativistic theories that accommodate long-range forces such as are observed in nature.
Let|p,±j∝an}b∇acket∇i}htand|p′,±j∝an}b∇acket∇i}htbe two one-particle, massless states of spin j, labelled by their
light-like 4-momenta pandp′, and by their helicity (which we take to be the same for the
two particles). We will be considering the matrix elements
/angbracketleftbig
p′,±j/vextendsingle/vextendsinglejµ|p,±j∝an}b∇acket∇i}ht;/angbracketleftbig
p′,±j/vextendsingle/vextendsingleTµν|p,±j∝an}b∇acket∇i}ht, (2.69)
wherejµis a conserved current (i.e., ∂µ∝an}b∇acketle{tjµ∝an}b∇acket∇i}ht= 0) andTµνis a conserved stress-energy
tensor (i.e. ∂µ∝an}b∇acketle{tTµν∝an}b∇acket∇i}ht= 0).
2.6.1 The j >1/2case
If we assume that the massless particles in question carry a n on-zero conserved charge
Q=/integraltext
d3xJ0, so that (suppressing the helicity label for now)
Q|p∝an}b∇acket∇i}ht=q|p∝an}b∇acket∇i}ht, (2.70)
9According to the authors, a less general version of their the orem was formulated earlier by Sidney
Coleman, but was not published.
24
whereq∝ne}ationslash= 0, then evidently
/angbracketleftbig
p′/vextendsingle/vextendsingleQ/vextendsingle/vextendsinglep/angbracketrightbig
=qδ3(p′−p). (2.71)
Meanwhile, we also have that
/angbracketleftbig
p′/vextendsingle/vextendsingleQ/vextendsingle/vextendsinglep/angbracketrightbig
=/integraldisplay
d3x/angbracketleftbig
p′/vextendsingle/vextendsinglej0(t,x)/vextendsingle/vextendsinglep/angbracketrightbig
=/integraldisplay
d3x/angbracketleftbig
p′/vextendsingle/vextendsingleeiP·xj0(t,0)e−iP·x/vextendsingle/vextendsinglep/angbracketrightbig
=/integraldisplay
d3xei(p′−p)·x/angbracketleftbig
p′/vextendsingle/vextendsinglej0(t,0)/vextendsingle/vextendsinglep/angbracketrightbig
= (2π)3δ3(p′−p)/angbracketleftbig
p′/vextendsingle/vextendsinglej0(t,0)/vextendsingle/vextendsinglep/angbracketrightbig
,(2.72)
so that combining Eqs. (2.71) and (2.72) gives
lim
p′→p/angbracketleftbig
p′/vextendsingle/vextendsinglej0(t,0)/vextendsingle/vextendsinglep/angbracketrightbig
=q
(2π)3, (2.73)
which, by Lorentz covariance, implies that
lim
p′→p/angbracketleftbig
p′/vextendsingle/vextendsinglejµ(t,0)/vextendsingle/vextendsinglep/angbracketrightbig
=qpµ
E(2π)3∝ne}ationslash= 0. (2.74)
Notice that Eq. (2.74) implies current conservation, becau sep2= 0.
For any light-like pandp′
(p′+p)2= 2(p′·p) = 2(|p′||p| −p′·p) = 2|p′||p|(1−cosθ)≥0, (2.75)
whereθis the angle between the momenta. If θ∝ne}ationslash= 0, then ( p′+p) is time-like and we can
therefore choose a frame in which it has no space component, s o that
p= (|p|,p) ;p′= (|p|,−p) (2.76)
(i.e., the two particles propagate in opposite directions w ith the same energy). In this frame,
consider rotating the particles by an angle φaround the axis of p:
|p,±j∝an}b∇acket∇i}ht →e±iφj|p,±j∝an}b∇acket∇i}ht;/vextendsingle/vextendsinglep′,±j/angbracketrightbig
→e∓iφj|p,±j∝an}b∇acket∇i}ht. (2.77)
25
The Lorentz covariance of the matrix element of jµthen implies that
e±2iφj/angbracketleftbig
p′,±j/vextendsingle/vextendsinglejµ(t,0)/vextendsingle/vextendsinglep,±j/angbracketrightbig
= Λ(φ)µ
ν/angbracketleftbig
p′,±j/vextendsingle/vextendsinglejν(t,0)/vextendsingle/vextendsinglep,±j/angbracketrightbig
, (2.78)
where Λ(φ) is the Lorentz transformation corresponding to a rotation by an angle φaround
the direction of p. But Λ(φ) contains no Fourier components other than e±iφand 1, so
Eq. (2.78) implies that the matrix elements vanish for j >1/2. In the limit p′→p,
we then arrive at a contradiction with Eq. (2.74). Therefore no relativistic QFT with a
conserved current can have massless spin-1 particles (eith er fundamental or composite) that
have Lorentz-covariant spectra and are charged under the co nserved current.
2.6.2 The j >1case
If the massless particles in question carry no conserved cha rge, we may still consider the
matrix elements of the stress-energy tensor Tµν. By the same kind of argument as in
Subsection 2.6.1
lim
p′→p/angbracketleftbig
p′/vextendsingle/vextendsingleTµν(t,0)/vextendsingle/vextendsinglep/angbracketrightbig
=pµpν
E(2π)3∝ne}ationslash= 0. (2.79)
Notice again that this stress-energy is conserved because p2= 0.
Then combining Eq. (2.77) with relativistic covariance imp lies that
e±2iφj/angbracketleftbig
p′,±j/vextendsingle/vextendsingleTµν(t,0)/vextendsingle/vextendsinglep,±j/angbracketrightbig
= Λ(φ)µ
ρΛ(φ)ν
σ/angbracketleftbig
p′,±j/vextendsingle/vextendsingleTρσ(t,0)/vextendsingle/vextendsinglep,±j/angbracketrightbig
. (2.80)
The fact that Λ( φ) contains only the Fourier components e±iφand 1 then implies that
the matrix elements must vanish for j >1, contradicting Eq. (2.79) in the limit p′→p.
Therefore no relativistic QFT with a conserved stress-ener gy tensor can have massless spin-2
particles (either fundamental or composite) that have Lore ntz-covariant spectra.
2.6.3 Why are gluons and gravitons allowed?
Evidently, the Weinberg-Witten theorem does not forbid pho tons, because they carry no
conserved charge. It also does not forbid the W±andZbosons because they are massive.
But the Standard Model contains charged, massless spin-1 pa rticles (the gluons) as well
as massless spin-2 particles (the gravitons). How is this po ssible? The resolution of this
question helps to clarify the necessity for local gauge inva riance.
26
In a Yang-Mills theory,
LYM=−1
4Fa
µνFaµν+Lmatter(ψ,D µψ) (2.81)
the gauge-invariant current
jµ
a=δSmatter
δAaµ(2.82)
is not conserved, because it obeys the equation Dµ∝an}b∇acketle{tjµ
a∝an}b∇acket∇i}ht= 0, rather than ∂µ∝an}b∇acketle{tjµ
a∝an}b∇acket∇i}ht= 0. Fur-
thermore, ∝an}b∇acketle{tjµ
a∝an}b∇acket∇i}htvanishes for one-particle gauge field states. Therefore con sidering the matrix
elements of this jµ
abetween gauge boson states in Yang-Mills theory would avail us nothing
because the limit in Eq. (2.74) would be zero.
What we actually want is a current that measures the flow of cha rge in the absence of
matter (i.e., for the Yang-Mills bosons alone) and that is co nserved in the sense ∂µ∝an}b∇acketle{tJµ
a∝an}b∇acket∇i}ht= 0:
Jµ
a=−Fµν
cfcabAbν, (2.83)
where thef’s are the structure constants of the gauge group. Conservat ion follows imme-
diately from the equation of motion for Eq. (2.81). This is, i n fact, the conserved current
obtained through Noether’s theorem from the global gauge in variance of Eq. (2.81) without
matter. But the current in Eq. (2.83) is obviously not gauge- invariant. Therefore, under
the action of a Lorentz transformation Λ,
Jµ
a→Λµ
νJν
a+∂µΩa (2.84)
and it is not, consequently, Lorentz-covariant. If we tried making it Lorentz-covariant by
introducing an unphysical extra polarization of the gauge b oson, then the theorem would
fail because the helicities would not be Lorentz-invariant , invalidating the choice-of-frame
procedure used to arrive at Eq. (2.77).
To put this in another way, in a gauge theory the physical |p,±j∝an}b∇acket∇i}htstates are actually
equivalence classes, because two states related by a gauge t ransformation represent the same
physics. A technical way of thinking about this is that the ph ysical states are elements of
the BRST cohomology ([13]). Therefore, matrix elements suc h as those in Eq. (2.69)
are only well-defined if the operator jµis BRST-closed, which requires the operator to be
27
gauge-invariant. It is well known that Yang-Mills theories do not allow the construction of
gauge-invariant conserved currents.
The case of the graviton is very closely analogous to that of t he Yang-Mills bosons. In
Einstein-Hilbert gravity
S=/integraldisplay
d4x√−g[R+Lmatter(φ,∇µφ,gµν)], (2.85)
where the field φstands for all possible matter fields of any spin. The covaria nt stress-energy
tensor
Tµν=1√−gδSmatter
δgµν(2.86)
obeys ∇µ∝an}b∇acketle{tTµν∝an}b∇acket∇i}ht= 0 rather than ∂µ∝an}b∇acketle{tTµν∝an}b∇acket∇i}ht= 0, and ∝an}b∇acketle{tTµν∝an}b∇acket∇i}ht= 0 for any state with only gravi-
tational fields. What we want is therefore not T, but rather
Θµ
ν=∂R
∂(∂µgαβ)(∂νgαβ)−gµ
νR . (2.87)
But recall that the Ricci scalar Rcontains not only the metric and its first derivatives, but
also terms linear in its second derivatives. In order to defin e Θ we therefore need to do the
usual trick of integrating by parts and setting the boundary terms to zero in order to get
rid of the second derivatives in R. This means that Ris no longer a covariant scalar and
therefore Θ is not a covariant tensor, but rather a pseudoten sor.
It is well known that gravitational energy cannot be defined i n a covariant way. For
instance, the energy of gravity waves on a flat background is l ocalizable only for waves trav-
eling in a single direction, which is not a coordinate-invar iant condition (see, for instance,
Chapter 33 in [14]). A general Lorentz transformation of the graviton field hµνwill destroy
this condition. This means that the stress-energy pseudote nsor Θµνfor gravitons involves
a fieldhµνthat does not transform like a Lorentz tensor. Its matrix ele ments are therefore
not Lorentz-covariant. Once again, if we attempt to remedy t his by introducing unphysical
extra polarizations of the gravitons, the Lorentz invarian ce of the helicity is lost.
Otherwise stated, in a theory with diffeomorphism invarianc e like GR, the physical
states are equivalence classes, because two states related by a coordinate transformation
represent the same physics. The matrix elements in Eq. (2.69 ) are only well defined if the
operatorTµνis BRST-closed, but GR admits no local BRST-closed operator s, and thus
28
evades the Weinberg-Witten theorem.
Notice that even in theories with a local symmetry, such as QC D or GR, the Weinberg-
Witten theorem does rule massless particles of higher spin t hat carry a conserved charge
associated with a symmetry that commutes with the local symm etry. For instance, the au-
thors of [4] point out that their result forbids QCD from havi ng flavor non-singlet massless
bound states with j≥1, since flavor symmetries commute with the SU(3) local gauge
symmetry. Similarly, a j= 1 gauge theory cannot produce composite gravitons with
Lorentz-covariant spectra, because translations in flat Mi nkowski space-time commute with
the gauge symmetry. Gauge theories admit the conserved, Lor entz-covariant Belinfante-
Rosenfeld stress-energy tensor ([15]).
2.6.4 Gravitons in string theory
String theories have a massless spin-2 particle in their spe ctrum. This discovery killed the
original versions of string theory as possible description s of the strong nuclear interaction
(which was the context in which they had been proposed) and ma de modern string theory a
candidate for a quantum theory of gravity (see, for instance , Chapter 1 in [16]). The reason
why this result does not violate the Weinberg-Witten theore m is that it is not possible to
define a conserved stress-energy tensor in string theory.
Consider a string propagating in a D-dimensional background space-time with metric
gab, wherea,b= 0,1,...D −1. IfSis the action in the background, then
Tab=1√−gδS
δgab(2.88)
is not well defined because a consistent string theory requir es imposing superconformal sym-
metry on the background, which in turn automatically requir esgabto obey an equation of
motion (at low energies this equation of motion corresponds to the Einstein field equation
of GR). The functional derivative in Eq.(2.88) cannot be defi ned because there is no con-
sistent off-shell definition of the background action S: The exact equation of motion for gab
in string theory does not come from extremizing the action wi th respect to the background
metric, but rather from a constraint required for consisten cy.10
In general, we expect that a theory with emergent diffeomorphism invariance would
10I thank John Schwarz for clarifying this point for me.
29
not have a stress-energy tensor. The reason is that in the low -energy effective action (i.e.,
in GR) the graviton couples to a stress-energy tensor which i s not observable because it
is not diffeomorphism-covariant. If the fundamental theory itself has no diffeomorphism
invariance, then it should not have a stress-energy tensor a t all (see [17]).
2.7 Emergent gravity
The Weinberg-Witten theorem can be read as the proof that mas sless particles of higher
spin cannot carry conserved Lorentz-covariant quantities . Local gauge invariance and dif-
feomorphism invariance are natural ways of making those qua ntities mathematically non-
Lorentz-covariant without spoiling physical Lorentz cova riance. It is possible and interesting
nonetheless, to consider other ways of accommodating massl ess mediators with higher spin.
Despite the successes of gauge theories, the fact remains th at there is no clearly compelling
a priori reason to impose local gauge invariance as an axiom, and that such an axiom has
the unattractive consequence that it makes our mathematica l description of physical reality
inherently redundant (see, for instance, Chapter III.4 in [ 18]).
Also, while local gauge invariance guarantees renormaliza bility for spin 1, it is well known
that quantizing hµνin linear gravity does not produce a perturbatively renorma lizable
field theory. One attractive solution to this problem would b e to make the graviton a
composite, low-energy degree of freedom, with a natural cut off scale Λ UV. The Weinberg-
Witten theorem represents a significant obstruction to this approach, because the result
applies equally to fundamental and to composite particles. Indeed, ruling out emergent
gravitons was the authors’ purpose for establishing that th eorem.
In a recent public lecture ([19]), Witten has made the strong claim that “whatever
we do, we are not going to start with a conventional theory of n on-gravitational fields in
Minkowski spacetime and generate Einstein gravity as an eme rgent phenomenon.” His
reasoning is that identifying emergent phenomena requires first defining a box in 3-space
and then integrating out modes with wavelengths shorter tha n the length of the edges of
the box (see Fig. 2.2). But Einstein gravity implies diffeomo rphism invariance, and a
general coordinate transformation spoils the definition of our box. Witten’s conclusion is
that gravity can be emergent only if the notion on the space-t ime on which diffeomorphism
invariance operates is simultaneously emergent. This is a p lausible claim, but it goes beyond
30
b^, 89[m [
Figure 2.2: Schematic representation of Witten’s argument that a gener al coordinate transformation
spoils the box used to define the modes that are integrated out in order to identify the emergent
low-energy physics for energy scales well below Λ UV.
what the Weinberg-Witten theorem actually establishes.
In 1983, Laughlin explained the observed fractional quantu m Hall effect in two-dimensional
electronic systems by showing how such a system could form an incompressible quantum
fluid whose excitations have charge e/3 ([20]). That is, the low-energy theory of the inter-
acting electrons in two spatial dimensions has composite de grees of freedom whose charge
is a fraction of that of the electrons themselves. In 2001, Zh ang and Hu used techniques
similar to Laughlin’s to study the composite excitations of a higher-dimensional system
([21]). They imagined a four-dimensional sphere in space, fi lled with fermions that interact
via anSU(2) gauge field. In the limit where the dimensionality of the r epresentation of
SU(2) is taken to be very large, such a theory exhibits composit e massless excitations of
integer spin 1, 2 and higher.
Like other theories from solid state physics, Zhang and Hu’s proposal falls outside the
scope of the Weinberg-Witten theorem because the proposed t heory is not Lorentz-invariant:
The vacuum of the theory is not empty and has a preferred rest- frame (the rest frame of
the fermions). However, the authors argued that in the three -dimensional boundary of
the four-dimensional sphere, a relativistic dispersion re lation would hold. One might then
imagine that the relativistic, three-dimensional world we inhabit might be the edge of a
four-dimensional sphere filled with fermions. Photons and g ravitons would be composite
low-energy degrees of freedom, and the problems currently a ssociated with gravity in the
UV would be avoided. The authors also argue that massless bos ons with spin 3 and higher
might naturally decouple from other matter, thus explainin g why they are not observed in
nature.
31
In Chapter 3 we will discuss another proposal, dating back to the work of Dirac ([22]) and
Bjorken ([23]) for obtaining massless mediators as the Gold stone bosons of the spontaneous
breaking of Lorentz violation. Such an arrangement evades t he Weinberg-Witten theorem
because the Lorentz invariance of the theory is realized non -linearly in the Goldstone bosons.
Therefore the matrix elements in Eq. (2.69) will not be Loren tz-covariant.
32
Chapter 3
Goldstone photons and gravitons
In this chapter we will address some issues connected with th e construction of models in
which massless mediators are obtained as Goldstone bosons o f the spontaneous breaking
of Lorentz invariance (LI). This presentation is based larg ely on previously published work
[24, 25].
3.1 Emergent mediators
In 1963, Bjorken proposed a mechanism for what he called the “ dynamical generation of
quantum electrodynamics” (QED) ([23]). His idea was to form ulate a theory that would
reproduce the phenomenology of standard QED, without invok ing localU(1) gauge invari-
ance as an axiom. Instead, Bjorken proposed working with a se lf-interacting fermion field
theory of the form
L=¯ψ(i∂ /−m)ψ−λ(¯ψγµψ)2. (3.1)
Bjorken then argued that in a theory such as that described by Eq. (3.1), composite
“photons” could emerge as Goldstone bosons resulting from t he presence of a condensate
that spontaneously broke LI.
Conceptually, a useful way of understanding Bjorken’s prop osal is to think of it as as a
resurrection of the “lumineferous æther” ([26, 27]): “empt y” space is no longer really empty.
Instead, the theory has a non-vanishing vacuum expectation value (VEV) for the current
jµ=¯ψγµψ. This VEV, in turn, leads to a massive background gauge field Aµ∝jµ, as in the
well-known London equations for the theory of superconduct ors ([28]). Such a background
spontaneously breaks Lorentz invariance and produces thre e massless excitations of Aµ(the
33
Goldstone bosons) proportional to the changes δjµassociated with the three broken Lorentz
transformations.1
Two of these Goldstone bosons can be interpreted as the usual transverse photons.
The meaning of the third photon remains problematic. Bjorke n originally interpreted it
as the longitudinal photon in the temporal-gauge QED, which becomes identified with the
Coulomb force (see also [26]). More recently, Kraus and Tomb oulis have argued that the
extra photon has an exotic dispersion relation and that its c oupling to matter should be
suppressed ([30]).
Bjorken’s idea might not seem attractive today, since a theo ry such as Eq. (3.1) is
not renormalizable, while the work of ’t Hooft and others has demonstrated that a lo-
cally gauge-invariant theory can always be renormalized ([ 5]). Furthermore, as detailed in
Section 2.4, the gauge theories have had other very significa nt successes. Unless we take
seriously the line of thought pursued in Chapter 2 that local gauge invariance is suspect
because it is a redundancy of the mathematical description r ather than a genuine physical
symmetry, there would not appear to be, at this stage in our un derstanding of fundamental
physics, any compelling reason to abandon local gauge invar iance as an axiom for writing
down interacting QFT’s.2Furthermore, the arguments for the existence of a LI-breaki ng
condensate in theories such as Eq. (3.1) have never been soli d.3
In 2002 Kraus and Tomboulis resurrected Bjorken’s idea for a different purpose of greater
interest to contemporary theoretical physics: making a com posite graviton ([30]). They
proposed what Bjorken might call “dynamical generation of g ravity.” In this scenario a
composite graviton would emerge as a Goldstone boson from th e spontaneous breaking of
Lorentz invariance in a theory of self-interacting fermion s. Being a Goldstone boson, such
a graviton would be forbidden from developing a potential, t hus providing a solution to the
“large cosmological constant problem:” the Λ hµ
µtadpole term for the graviton would vanish
without fine-tuning (see Section 5.1). This scheme would als o seem to offer an unorthodox
avenue to a renormalizable quantum theory of gravity, becau se the fermion self-interactions
1In Bjorken’s work, Aµis just an auxiliary or interpolating field. Dirac had discus sed somewhat similar
ideas in [22], but, amusingly, he was trying to write a theory of electromagnetism with only a gauge field and
no fundamental electrons. In both the work of Bjorken and the work of Dirac, the proportionality between
Aµandjµis crucial.
2According to Mark Wise, though, in the 1980’s Feynman consid ered Bjorken’s proposal as an alternative
to postulating local gauge invariance.
3For Bjorken’s most recent revisiting of his proposal, in the light of the theoretical developments since
1963, see [29].
34
could be interpreted as coming from the integrating out, at l ow energies, of gauge bosons
that have acquired large masses via the Higgs mechanism, so t hat Einstein gravity would
be the low energy behavior of a renormalizable theory. This p roposal would, of course,
radically alter the nature of gravitational physics at very high energies. Related ideas had
been previously considered in, for instance, [31].
In [30], the authors consider fermions coupled to gauge boso ns that have acquired masses
beyond the energy scale of interest. Then an effective low-en ergy theory can be obtained
by integrating out those gauge bosons. We expect to obtain an effective Lagrangian of the
form
L=¯ψ(i∂ /−m)ψ+∞/summationdisplay
n=1λn(¯ψγµψ)2n
+∞/summationdisplay
n=1µn/bracketleftbigg
¯ψi
2(γµ→
∂ν−γµ←
∂ν)ψ/bracketrightbigg2n
+... , (3.2)
where we have explicitly written out only two of the power ser ies in fermion bilinears that
we would in general expect to get from integrating out the gau ge bosons.
One may then introduce an auxiliary field for each of these fer mion bilinears. In this
example we shall assign the label Aµto the auxiliary field corresponding to ¯ψγµψ, and
the labelhµνto the field corresponding to ¯ψi
2(γµ→
∂ν−γµ←
∂ν)ψ. It is possible to write
a Lagrangian that involves the auxiliary fields but not their derivatives, so that the cor-
responding algebraic equations of motion relating each aux iliary field to its corresponding
fermion bilinear make that Lagrangian classically equival ent to Eq. (3.2). In this case the
new Lagrangian would be of the form
L′= (ηµν+hµν)¯ψi
2(γµ→
∂ν−γµ←
∂ν)ψ−¯ψ(A /+m)ψ+...
−VA(A2)−Vh(h2) +... , (3.3)
whereA2≡AµAµandh2≡hµνhµν. The ellipses in Eq. (3.3) correspond to terms with
other auxiliary fields associated with more complicated fer mion bilinears that were also
omitted in Eq. (3.2).
We may then imagine that instead of having a single fermion sp ecies we have one very
heavy fermion, ψ1, and one lighter one, ψ2. Since Eq. (3.3) has terms that couple both
35
fermion species to the auxiliary fields, integrating out ψ1will then produce kinetic terms
forAµandhµν.
In the case of Aµwe can readily see that since it is minimally coupled to ψ1, the kinetic
terms obtained from integrating out the latter must be gauge -invariant (provided a gauge-
invariant regulator is used). To lowest order in derivative s ofAµ, we must then get the
standard photon Lagrangian −F2
µν/4. SinceAµwas also minimally coupled to ψ2, we then
have, at low energies, something that has begun to look like Q ED.
IfAµhas a non-zero VEV, LI is spontaneously broken, producing th ree massless Gold-
stone bosons, two of which may be interpreted as photons (see [30] for a discussion of how
the exotic physics of the other extraneous “photon” can be su ppressed). The integrating
out ofψ1and the assumption that hµνhas a VEV, by similar arguments, yield a low-energy
approximation to linearized gravity.
Fermion bilinears other than those we have written out expli citly in Eq. (3.2) have their
own auxiliary fields with their own potentials. If those pote ntials do not themselves produce
VEV’s for the auxiliary fields, then there would be no further Goldstone bosons, and one
would expect, on general grounds, that those extra auxiliar y fields would acquire masses of
the order of the energy-momentum cutoff scale for our effectiv e field theory, making them
irrelevant at low energies.
The breaking of LI would be crucial for this kind of mechanism , not only because we
know experimentally that photons and gravitons are massles s or very nearly massless, but
also because it allows us to evade the Weinberg-Witten theor em ([4]), as we discussed in
Section 2.7.
Let us concentrate on the simpler case of the auxiliary field Aµ. For the theory described
by Eq. (3.3), the equation of motion for Aµis
∂L′
∂Aµ=−¯ψγµψ−V′(A2)·2Aµ= 0. (3.4)
Solving for ¯ψγµψin Eq. (3.4) and substituting into both Eq. (3.2) and Eq. (3.3 ) we see
that the condition for the Lagrangians LandL′to be classically equivalent is a differential
equation for V(A2) in terms of the coefficients λn:
V(A2) = 2A2[V′(A2)]−∞/summationdisplay
n=1λn22nA2n[V′(A2)]2n. (3.5)
36
It is suggested in [30] that for some values of λnthe resulting potential V(A2) might
have a minimum away from A2= 0, and that this would give the LI-breaking VEV needed.
It seems to us, however, that a minimum of V(A2) away from the origin is not the correct
thing to look for in order to obtain LI breaking. The Lagrangi an in Eq. (3.3) contains
Aµ’s not just in the potential but also in the “interaction” ter mAµ¯ψγµψ, which is not in
any sense a small perturbation as it might be, say, in QED. In o ther words, the classical
quantityV(A2) is not a useful approximation to the quantum effective poten tial for the
auxiliary field.
In fact, regardless of the values of the λn, Eq. (3.5) implies that V(A2= 0) = 0, and also
that at any point where V′(A2) = 0 the potential must be zero. Therefore, the existence
of a classical extremum at A2=C∝ne}ationslash= 0 would imply that V(C) =V(0), and unless the
potential is discontinuous somewhere, this would require t hatV′(and therefore also V)
vanish somewhere between 0 and C, and so on ad infinitum . Thus the potential Vcannot
have a classical minimum away from A2= 0, unless the potential has poles or some other
discontinuity.
A similar observation applies to any fermion bilinear for wh ich we might attempt this
kind of procedure and therefore the issue arises as well when dealing with the proposal in
[30] for generating the graviton. It is not possible to sides tep this difficulty by including
other auxiliary fields or other fermion bilinears, or even by imagining that we could start,
instead of from Eq. (3.2), from a theory with interactions gi ven by an arbitrary, possibly
non-analytic function of the fermion bilinear F(bilinear). The problem can be traced to
the fact that the equation of motion of any auxiliary field of t his kind will always be of the
form
0 =−(bilinear) −V′(field2)·2field. (3.6)
The point is that the vanishing of the first derivative of the p otential or the vanishing
of the auxiliary field itself will always, classically, impl y that the fermion bilinear is zero.
Classically at least, it would seem that the extrema of the po tential would correspond to
the same physical state as the zeroes of the auxiliary field.
37
3.2 Nambu and Jona-Lasinio model (review)
The complications we have discussed that emerge when one tri es to implement LI breaking
as proposed in [30] do not, in retrospect, seem entirely surp rising. A VEV for the auxiliary
field would classically imply a VEV for the corresponding fer mion bilinear, and therefore a
trick such as rewriting a theory in a form like Eq. (3.3) shoul d not, perhaps, be expected
to uncover a physically significant phenomenon such as the sp ontaneous breaking of LI for
a theory where it was not otherwise apparent that the fermion bilinear in question had a
VEV. Let us therefore turn our attention to considering what would be required so that
one might reasonably expect a fermion field theory to exhibit the kind of condensation that
would give a VEV to a certain fermion bilinear.
If we allowed ourselves to be guided by purely classical intu ition, it would seem likely
that a VEV for a bilinear with derivatives (such as ¯ψi
2(γµ→
∂ν−γµ←
∂ν)ψ) might require non-
standard kinetic terms in the action. Whether or not this int uition is correct, we abandon
consideration of such bilinears here as too complicated.
The simplest fermion bilinear is, of course, ¯ψψ. Being a Lorentz scalar, ∝an}b∇acketle{t¯ψψ∝an}b∇acket∇i}ht ∝ne}ationslash= 0 will
not break LI. This kind of VEV was treated back in 1961 by Nambu and Jona-Lasinio,
who used it to spontaneously break chiral symmetry in one of t he early efforts to develop
a theory of the strong nuclear interactions, before the adve nt of quantum chromodynamics
(QCD) ([32]). It might be useful to review the original work o f Nambu and Jona-Lasinio,
as it may shed some light on the study of the possibility of giv ing VEV’s to other fermion
bilinears that are not Lorentz scalars.
In their original paper, Nambu and Jona-Lasinio start from a self-interacting massless
fermion field theory and propose that the strong interaction s be mediated by pions, which
appear as Goldstone bosons produced by the spontaneous brea king of the chiral symmetry
associated with the transformation ψ∝ma√sto→exp (iαγ5)ψ. This symmetry breaking is produced
by a VEV for the fermion bilinear ¯ψψ. In other words, Nambu and Jona-Lasinio originally
proposed what, by close analogy to Bjorken’s idea, would be t he “dynamical generation of
the strong interactions.”4
Nambu and Jona-Lasinio start from a non-renormalizable qua ntum field theory with a
4Historically, though, Bjorken was motivated by the earlier work of Nambu and Jona-Lasinio.
38/A1
/BP/A2
/B7/A3
/BD/C8/C1
/BC/BP/A4
/B7/A5
/BD/C8/C1
/BC/B7/A6
/BD/C8/C1
/BC/BD/C8/C1
/BC/B7 /BM /BM /BM
Figure 3.1: Diagrammatic Schwinger-Dyson equation. The double line re presents the primed prop-
agator, which incorporates the self-energy term. The singl e line represents the unprimed propagator.
1PI′stands for the sum of one-particle irreducible graphs with t he primed propagator.
four-fermion interaction that respects chiral symmetry:
L=i¯ψ∂ /ψ−g
2[(¯ψγµψ)2−(¯ψγµγ5ψ)2]. (3.7)
In order to argue for the presence of a chiral symmetry-break ing condensate in the
theory described by Eq. (3.7), Nambu and Jona-Lasinio borro wed the technique of self-
consistent field theory from solid state physics (see, for in stance, [33]). If one writes down
a Lagrangian with a free and an interaction part, L=L0+Li, ordinarily one would then
proceed to diagonalize L0and treat Lias a perturbation. In self-consistent field theory one
instead rewrites the Lagrangian as L= (L0+Ls) + (Li− Ls) =L′
0+L′
i, where Lsis a
self-interaction term, either bilinear or quadratic in the fields, such that L′
0yields a linear
equation of motion. Now L′
0is diagonalized and L′
iis treated as a perturbation.
In order to determine what the form of Lsis, one requires that the perturbation L′
inot
produce any additional self-energy effects. The name “self- consistent field theory” reflects
the fact that in this technique Liis found by computing a self-energy via a perturbative
expansion in fields that already are subject to that self-ene rgy, and then requiring that such
a perturbative expansion not yield any additional self-ene rgy effects.
Nambu and Jona-Lasinio proceed to make the ansatz that for Eq . (3.7) the self-
interaction term will be of the form Ls=−m¯ψψ. Then, to first order in the coupling
constantg, they proceed to compute the fermion self-energy Σ′(p), using the propagator
S′(p) =i(p /−m)−1, which corresponds to the Lagrangian L′
0=¯ψ(i∂ /−m)ψthat incorporates
the proposed self-energy term.
The next step is to apply the self-consistency condition usi ng the Schwinger-Dyson
39/A0 /CX /A6
/BC/BP/A1
/BD/C8/C1
/BC/BP/A2
/B7 /C7 /B4 /CV
/BE/B5
Figure 3.2: Diagrammatic equation for the primed self-energy. We will w ork to first order in the
fermion self-coupling constant g.
equation for the propagator
S′(x−y) =S(x−y) +/integraldisplay
d4zS(x−z)Σ′(0)S′(z−y), (3.8)
which is represented diagrammatically in Fig. 3.1. The prim es indicate quantities that
correspond to a free Lagrangian L′
0that incorporates the self-energy term, whereas the
unprimed quantities correspond to the ordinary free Lagran gianL0. For Σ′we will use the
approximation shown in Fig. 3.2, valid to first order in the co upling constant g.
After Fourier transforming Eq. (3.8) and summing the left si de as a geometric series,
we find that the self-consistency condition may be written, i n our approximation, as
m= Σ′(0) =gmi
2π4/integraldisplayd4p
p2−m2+iǫ. (3.9)
If we evaluate the momentum integral by Wick rotation and reg ularize its divergence
by introducing a Lorentz-invariant energy-momentum cutoff p2<Λ2we find
2π2m
gΛ2=m/bracketleftbigg
1−m2
Λ2log/parenleftbiggΛ2
m2+ 1/parenrightbigg/bracketrightbigg
. (3.10)
This equation will always have the trivial solution m= 0, which corresponds to the
vanishing of the proposed self-interaction term Li. But if
0<2π2
gΛ2<1 (3.11)
then there may also be a non-trivial solution to Eq. (3.10), i .e., a non-zero mfor which the
condition of self-consistency is met. For a rigorous treatm ent of the relation between non-
trivial solutions of this self-consistent equation and loc al extrema in the Wilsonian effective
potential for the corresponding fermion bilinears, see [39 ] and the references therein.
40
In this model (which from now on we shall refer to as NJL), we se e that if the interaction
between fermions and antifermions is attractive ( g >0) and strong enough (2π2
gΛ2<1) it
might be energetically favorable to form a fermion-antifer mion condensate. This is reason-
able to expect in this case because the particles have no bare mass and thus the energy
cost of producing them is small. The resulting condensate wo uld have zero net charge,
as well as zero total momentum and spin. Therefore it must pai r a left-handed fermion
ψL=1
2(1−γ5)ψwith the antiparticle of a right-handed fermion ψR=1
2(1+γ5)ψ, and vice
versa. This is the mass-term self-interaction Li=−m¯ψψ=−m(¯ψLψR+¯ψRψL) that NJL
studies.
After QCD became the accepted theory of the strong interacti ons, the ideas behind the
NJL mechanism remained useful. The uanddquarks are not massless (nor is u-dflavor
isospin an exact symmetry) but their bare masses are believe d to be quite small compared to
their effective masses in baryons and mesons, so that the form ation of ¯uuand¯ddcondensates
represents the spontaneous breaking of an approximate chir al symmetry. Interpreting the
pions (which are fairly light) as the pseudo-Goldstone boso ns generated by the spontaneous
breaking of the approximate SU(2)R×SU(2)Lchiral isospin symmetry down to just SU(2),
proved a fruitful line of thought from the point of view of the phenomenology of the strong
interaction.5
Condition Eq. (3.11) has a natural interpretation if we thin k of the interaction in Eq.
(3.7) as mediated by massive gauge bosons with zero momentum and coupling e. For it
to be reasonable to neglect boson momentum in the effective th eory, the mass µof the
bosons should be µ>Λ. Ife2<2π2theng=e2/µ2<2π2/Λ2, which violates Eq. (3.11).
Therefore for chiral symmetry breaking to happen, the coupl ingeshould be quite large,
making the renormalizable theory nonperturbative. This is acceptable because the factor
of 1/µ2allows the perturbative calculations we have carried out in the effective theory Eq.
(3.7). This is why the NJL mechanism is modernly thought of as a model for a phenomenon
of non-perturbative QCD.
5For a treatment of this subject, including a historical note on the influence of the NJL model in the
development of QCD, see Chap. 19, Sec. 4 in [38].
41
3.3 An NJL-style argument for breaking LI
We have reviewed how NJL formulated a model that exhibited a n on-zero VEV for the
fermion bilinear ¯ψψ. The next simplest fermion bilinear that we might consider i s¯ψγµψ,
which was the one that Bjorken, Kraus, and Tomboulis conside red when they discussed the
“dynamical generation of QED.” This particular fermion bil inear is especially interesting
because it corresponds to the U(1) conserved current, and also because it is the simplest
bilinear with an odd number of Lorentz tensor indices, so tha t a non-zero VEV for it would
break not only LI but also charge (C), charge-parity (CP), an d charge-parity-time (CPT)
reversal invariance. C and CP may not be symmetries of the Lag rangian, as indeed they
are not in the standard model, but by a celebrated result CPT m ust be an invariance of any
reasonable theory (see [41] and references therein). This i nvariance, however, may well be
spontaneously broken, as it would be by any VEV with an odd num ber of Lorentz indices.
Before proceeding, however, it may be advisable to try to dev elop some physical intuition
about what would be required for a fermion bilinear like ¯ψγµψto exhibit a VEV. If we
choose a representation of the gamma matrix algebra and use i t to write out ( ¯ψγµψ)2for
an arbitrary Dirac bispinor ψ, we may check that ( ¯ψγµψ)2≥0 for the choice of mostly
negative metric gµν= diag(1,−1,−1,−1). That is, ¯ψγµψis time-like. This has an intuitive
explanation, based on the observation that ¯ψγµψis a conserved fermion-number current
density. Classically a charge density ρmoving with a velocity vwill produce a current
jµ= (ρ,ρv) (in units of c= 1). Therefore the relativistic requirement that the charg e
density not move faster than the speed of light in any frame of reference implies that
j2≥0. Considerations of causality make it natural to expect tha t something similar would
be true of ¯ψγµψ.
For any time-like Lorentz vector nµit is possible to find a Lorentz transformation that
maps it to a vector n′µwith only one non-vanishing component: n′0. For a constant current
densityjµ, this means that for jµto be non-zero there must be a charge density j0, which
has a rest frame. Therefore we only expect to see a VEV for ¯ψγµψif our theory somehow has
a vacuum with a non-zero fermion number density. The consequ ent spontaneous breaking
of LI may be seen as the introduction of a preferred reference frame: the rest frame of the
vacuum charge.
In the literature of finite density quantum field theory and of color superconductivity
42
zero density finite density
particlehole (antiparticle)
E=0
Figure 3.3: Fermion and antifermion energies in QFT, at zero density (le ft) and at finite density
(right). Finite density introduces a chemical potential te rm−f·¯ψγ0ψinto the fermion Lagrangian.
(see, for instance, [34] and [35]), the Lagrangians discuss ed are explicitly non-Lorentz-
invariant because they contain chemical potential terms of the formf·¯ψγ0ψ. This term
appears in theories whose ground state has a non-zero fermio n number because, by the
Pauli exclusion principle, new fermions must be added just a bove the Fermi surface, i.e.,
at energies higher than those already occupied by the pre-ex isting fermions, while holes
(which can be thought of as antifermions) should be made by re moving fermions at that
Fermi surface. The result is an energy shift that depends on t he number of fermions already
present and which has opposite signs for fermions and antife rmions, as illustrated in Fig.
3.3.
The physical picture that emerges is now, hopefully, cleare r: A theory with a VEV for
¯ψγµψis one with a condensate that has non-zero fermion number. Th is means that only
theories with some form of attractive interaction between p articles with the same sign in
fermion number may be expected to produce such a VEV. The situ ation is closely analogous
to BCS superconductivity ([40]), in which a phonon-mediate d attractive interaction between
electrons allows the presence of a condensate with non-zero electric charge. Note that in
the NJL model, the condensate was composed of fermion-antif ermion pairs, and therefore
clearly ∝an}b∇acketle{t¯ψγ0ψ∝an}b∇acket∇i}ht= 0, which implies ∝an}b∇acketle{t¯ψγµψ∝an}b∇acket∇i}ht= 0. It should now be clear why a VEV for ¯ψγµψ
would break not only LI but also C, CP, and CPT. This picture al so helps to clarify the
nature of the Goldstone bosons that we will be invoking as med iators of the electromagnetic
interaction: They are density waves in the background “Dira c sea,” whose energy at infinite
wavelengths vanishes because they are then proportional to the broken boosts.
43
There is an easy way to write a theory that will have a VEV for a U(1) conserved
current: to couple a massive photon to such a current via a pur ely imaginary charge. To
see this, let us write a Proca Lagrangian for a massive photon field with an external source:
L=−1
4F2
µν+µ2
2A2−jµAµ. (3.12)
The equation of motion for the photon field is
∂µFµν=jν−µ2Aν. (3.13)
At energy scales well below the photon mass µ, the kinetic term −F2
µν/4 may be ne-
glected with respect to the mass term µ2A2/2. We may then integrate out the photon at
zero momentum by solving the equation of motion Eq. (3.13) fo r the photon field Aµwith
its conjugate momenta Fµνset to zero, and substituting the result back into the Lagran gian
in Eq. (3.12). The resulting low-energy effective field theor y has the Hamiltonian
Heffective =j2
2µ2. (3.14)
Nothing interesting happens if the source is a timelike curr ent density, since in that case
Eq. (3.14) has its minimum at jµ= 0. But if we were to make the charge coupling to the
photon imaginary (e.g., jµ=ie¯ψγµψforereal), then j2is actually always negative (recall
that ( ¯ψγµψ)2is always positive) and we get a “potential” with the wrong si gn, so that the
energy can be made arbitrarily low by decreasing j2. If we make jµdynamical by adding to
the Lagrangian terms corresponding to the field that sets up t he current, we might expect,
for certain parameters in the theory, that the energy be mini mized for a finite value of jµ.
By making the charge purely imaginary, our effective theory a t energy scales much
lower than the photon mass µwill look similar to Eq. (3.7), except that the four-fermion
interaction in the effective Lagrangian will be e2(¯ψγµψ)2/2µ2(with an overall positive,
rather than a negative, sign). What this means is that fermio ns are attracting fermions and
antifermions are attracting antifermions, rather than wha t we had in NJL (and in QED):
attraction between a fermion and an antifermion. Condensat ion, if it occurs, will here
produce a net fermion number, spontaneously breaking C, CP, and CPT.6
6Dyson argued that a theory with a long-range attraction betw een particles of the same fermion number
44/A1=/A2+/A3
Figure 3.4: The four-fermion vertex in the self-interacting theory may be seen as the sum of two
photon-mediated interactions with a massive photon that ca rries zero momentum and is coupled to
the fermion via a purely imaginary charge.
Let us analyze this situation again more rigorously using se lf-consistent field theory
methods, following Nambu and Jona-Lasinio. For this we cons ider a fermion field with the
usual free Lagrangian L0=¯ψ(i∂ /−m0)ψand pose as our self-consistent ansatz:
Ls=−(m−m0)¯ψψ−f¯ψγ0ψ. (3.15)
The corresponding momentum-space propagator for L′
0=L0+Lsis, therefore,
S′(k) =i(k /−fγ0−m)−1. (3.16)
Now let us suppose that the interaction term looks like
Li=g
2(¯ψγµψ)2. (3.17)
To obtain the Feynman rules corresponding to Eq. (3.17) we no te that this is what
we would obtain in massive QED if we replaced the charge ebyieand the usual photon
propagator by igµν/µ2, withg=e2/µ2. Therefore to compute the self-energy we will rely
on the identity represented in Fig. 3.4. (In QED the second di agram on the right-hand side
of Fig. 3.4 would vanish by Furry’s theorem, but in our case th e propagator in the loop
will have a chemical potential term that breaks the C invaria nce on which Furry’s theorem
depends.)
would be unstable and used this to suggest that perturbative series in QED would diverge after renormaliza-
tion of the charge and mass [42]. As we will see at the end of thi s section, the “photon” mass µwill prevent
the instability in our case.
45
To leading order in g, the self-energy is
Σ(0) = 2ig/integraldisplayd4k
(2π)43(k0−f)γ0+ 3kiγi−2m
k2
0−k2−m2+f2−2fk0+iǫσ(3.18)
whereσ(a function of |k|,f, andm) takes values ±1 so as to enforce the standard Feynman
prescription for shifting the k0poles: positive k0poles are shifted down from the real line,
while negative poles are shifted up.
At first sight it might appear as if the self-energy in Eq. (3.1 8) could not be used to
argue for the breaking of LI, because the shift in the integra tion variable k∝ma√sto→k′= (k0−f,k)
would wipe out fdependence. This, however, is not the case, as we will see. We may carry
out thedk0integration, for which we must find the corresponding poles. These are located
at
k0=f±/radicalbig
k2+m2. (3.19)
From now on, without loss of generality, we will take fto be positive. The contour
integral that results from closing the d0kintegral of Eq. (3.18) in the complex plane will
vanish unless f <√
k2+m2, because otherwise both poles in Eq. (3.19) will lie on the
same side of the imaginary axis. In light of the Feynman presc ription used for the shifting
of the poles away from the real axis, it would then be possible to close the contour at infinity
so that there would be no poles in the interior. The pole-shif ting prescription, through its
effect on the dk0integral, is what introduces an actual fdependence into the expression
for the self-energy.
By the Cauchy integral formula, we have
Σ(0) =−g
4π3/integraldisplay
d3k/bracketleftigg
3√
k2+m2γ0+ 2m
2√
k2+m2
×θ(/radicalbig
k2+m2−f)−3
2γ0/bracketrightbigg
, (3.20)
where the second term in the right-hand side subtracts the co ntribution from closing the
contour out at infinity in the complex plane (note the branch c ut in the logarithm that
results from computing that part of the contour integral exp licitly). We will introduce the
cutoff k2<Λ2to make the integral in Eq. (3.20) finite.7
7Carrying out the dk0integration separately from the spatial integral is legiti mate and useful in light of
the form of Eq. (3.18), which does not lend itself naturally t o Wick rotation. But the use of a non-Lorentz-
46
-2000 -1000 0 1000 2000-1000-500050010001500
m
(a)-2 -1 0 1 205101520
m
(b)-2000 -1000 0 1000 2000-1000-500050010001500
m
(c)
-2000 -1000 0 1000 2000-1000-500050010001500
m
(d)-2 -1 0 1 205101520
m
(e)-2 -1 0 1 205101520
m
(f)
Figure 3.5: Plots of the left-hand side (in gray) and right-hand side (in black) of equation Eq.
(3.25). Define α≡g
2π2. For each plot the parameters are: (a) Λ = 100, m0= 0,α= 0.001. (b)
Λ = 100,m0= 15,α= 0.001. (c) Λ = 100, m0= 1200,α= 0.001. (d) Λ = 100, m0= 0,α= 0.002.
(e) Λ = 100, m0= 15,α= 0.002. (f) Λ = 200, m0= 15,α= 0.001.
Note that the Heaviside step function θ(√
k2+m2−f) in Eq. (3.20) is always unity if
m>f , so that there will be no fdependence at all in Eq. (3.20) unless m≤f. Assuming
thatm≤fwe have
Σ(0) =−g
2π2/bracketleftig
−(f2−m2)3/2γ0+m3log (f+/radicalbig
f2−m2)
−m3log (Λ +/radicalbig
Λ2+m2)
+mΛ/radicalbig
Λ2+m2−mf/radicalbig
f2−m2/bracketrightbigg
. (3.21)
As before, we use the Schwinger-Dyson equation Eq. (3.8), an d after summing up the
invariant regulator may cause concern that any breaking of L I we might arrive at could be an artifact of
our choice of regulator. An alternative is to regulate Eq. (3 .20) dimensionally by replacing d3kwithdd−1k.
The resulting equations are more complicated and the depend ence on the range of energies where our non-
renormalizable theory is valid is obscured, but the overall argument does not change. It is also possible to
multiply the integrand in Eq. (3.18) by a cutoff in Minkowski s paceθ(Λ2+k2) =θ(Λ2+k2
0−k2). For k2<Λ2
we get the same result as in Eq. (3.20). For k2>Λ2we must impose the condition that k2
0>k2−Λ2.
It should be pointed out that previous work on LI breaking has used 3-momentum cutoffs in computing
self-energies [56], although in that case there seems to be a physical interpretation for such a cutoff which
does not apply to the present discussion. The original work o f Nambu and Jona-Lasinio [32] considers cutoffs
in Euclidean 4-momentum and in 3-momentum, arriving in both cases at similar conclusions.
47
right-hand side as a geometric series, we arrive at the self- consistency condition for our
ansatz Eq. (3.15):
m0−m−fγ0=−Σ(0)
=g
2π2/bracketleftig
−(f2−m2)3/2γ0
+m3log/parenleftigg
f+/radicalbig
f2−m2
Λ +√
Λ2+m2/parenrightigg
+mΛ/radicalbig
Λ2+m2
−mf/radicalbig
f2−m2/bracketrightbigg
. (3.22)
Clearly Eq. (3.22) will not admit a non-trivial solution f∝ne}ationslash= 0 unlessgis positive, which
agrees with our intuition that the theory must exhibit attra ction between particles of the
same fermion number. The self-consistent condition Eq. (3. 22) may be separated into two
simultaneous equations:
f=g
2π2(f2−m2)3/2(3.23)
and
m0−m=gm
2π2/bracketleftigg
m2log/parenleftigg
f+/radicalbig
f2−m2
Λ +√
Λ2+m2/parenrightigg
+ Λ/radicalbig
Λ2+m2−f/radicalbig
f2−m2/bracketrightbigg
. (3.24)
It is important to bear in mind that Eqs. (3.23) and (3.24) wer e written under the assump-
tion thatf≥m. Forf <m thefdependence of the self-energy in Eq. (3.18) disappears.
The trivial, Lorentz-invariant solution f= 0 to the self-consistent equations will always
be present for any m, as should be the case when spontaneous breaking of a symmetr y is
observed.
Equation (3.23) can be readily solved for fas a function of m(imposing the condition
thatfbe real and positive), and the resulting f(m) can be substituted into Eq. (3.24) to
48
yield
m0−m=gm
2π2/bracketleftigg
m2log/parenleftigg
f(m) +/radicalbig
f2(m)−m2
Λ +√
Λ2+m2/parenrightigg
+ Λ/radicalbig
Λ2+m2−f(m)/radicalbig
f2(m)−m2/bracketrightbigg
. (3.25)
Equation (3.25) cannot be solved algebraically, but we may s tudy some of its properties
graphically. In Fig. 3.5 we have plotted the left-hand side a nd the right-hand side of Eq.
(3.25) for various values of the parameters g,m0, and Λ. As plot (a) illustrates, m0= 0
impliesm= 0, i.e., we cannot dynamically generate both a chemical pot ential and a mass
term. Form=m0= 0 we have
f=π/radicalbig
2/g. (3.26)
Plot (b) in Fig. 3.5 shows a 0 < m 0≪Λ for which the corresponding mwill be
significantly less than m0. Plot (c) in the same figure illustrates that a very large m0is
needed before m>m 0, but such solutions are not physically meaningful because m0itself
is already well beyond the energy scale for which our effectiv e theory is supposed to hold.
By comparing plot (b) to plot (e) we may see the effect of increa singgfor a given m0and
Λ. A comparison of plots (b) and (f) should illustrate the effe ct of increasing Λ with the
other parameters fixed.
The plots in Fig. 3.6 illustrate the progression, as the para meter Λ is increased for
fixedα, from an unstable theory in which bare masses m0on the order of Λ are mapped to
m>Λ, to a theory that maps such bare masses to m<Λ. Such an analysis of Eq. (3.25)
reveals that the condition for this mass stability is
0<2π2
gΛ2<1, (3.27)
which is reminiscent of the condition Eq. (3.11) for chiral s ymmetry breaking in the NJL
model (except that now the interaction has the opposite sign ). Combining Eq. (3.27) with
Eq. (3.26) (which was exact for m0but may serve approximately for m0small) we arrive
at the requirement
0<f2<Λ2, (3.28)
which would surely have to hold if our theory were stable. Ind eed, we may interpret Eq.
49
-10 -5 0 5 10-7.5-5-2.502.557.510
m
(a)-20 -10 0 10 20-15-10-505101520
m
(b)
-20 -10 0 10 20-15-10-505101520
m
(c)-20 -10 0 10 20-15-10-505101520
m
(d)
Figure 3.6: Plots of the left-hand side (in gray) and right-hand side (in black) of equation Eq.
(3.25). For all of them α≡g
2π2= 0.01. (a) Λ = m0= 2. (b) Λ = m0= 8. (c) Λ = m0= 12. (d)
Λ =m0= 16.
(3.28) as saying that if we pick physically good parameters g,m0, and Λ we will have a
stable theory with finite chemical potential f. The parameters for plots (a), (b), (d), (e),
and (f) in Fig. 3.5 all give examples of such stable theories. As in NJL, the good parameters
involveg−1/2large with respect to Λ, suggesting that Eq. (3.17) should be a low-energy
approximation to a non-perturbative interaction of a full r enormalizable theory that allows
attraction between particles of the same fermion number sig n.
The issue of how the form of the self-consistent equations wi ll depend on the choice of
regulator for the integral in Eq. (3.18) is not an entirely st raightforward matter. But it
seems to be a solid conclusion that, for positive fermion sel f-couplingg, the solutions to
such self-consistent equations show the presence of LI-bre aking vacua. In the next section
of this paper we offer an alternative approach that strengthe ns this conclusion and that
sheds further light on the issue of stability.
50
Γ[A] =V(A) +/A1+/A2+/A3+...
Figure 3.7: Correction of the effective potential of the auxiliary field Aµfrom integrating out the
fermion. The first graph does not contribute by the Ward ident ity, while the second vanishes by
Furry’s theorem.
3.4 Consequences for emergent photons
The theory
L=¯ψ(i∂ /−m0)ψ+g
2(¯ψγµψ)2(3.29)
is equivalent to
L′=¯ψ(i∂ /−A /−m0)ψ−A2
2g. (3.30)
Since we argued that Eq. (3.29) may spontaneously break LI by giving a finite ∝an}b∇acketle{t¯ψγµψ∝an}b∇acket∇i}ht,
we conclude that Aµin Eq. (3.30) would also have a finite VEV, since, by the algebr aic
equation of motion,
Aµ=−g¯ψγµψ. (3.31)
This interpretation agrees with the observation that Eq. (3 .30) has a vector boson field
whose mass term carries the wrong sign if g >0, indicating that the zero-field state is not
a good vacuum. To find the correct vacuum for the theory we must carry out the path
integral over the fermion field to obtain the effective action Γ[A], and then minimize that
quantity. Figure 3.7 shows the radiative corrections to Γ[ A] as a perturbative series, in terms
of Feynman diagrams. The field Aµis minimally coupled to ψ, so that the computation
should proceed as in QED. By the Ward identity we do not expect a correction to the mass
term forAµ, as long as an adequate regulator is used. But we do expect to g et terms in the
effective action that go as A4and higher even powers of the auxiliary field.
Since we have reason to believe that QED is stable for any valu e of the charge e, it
therefore seems logical to expect that the effective action f orAµin Eq. (3.30) gives it a
finite time-like VEV, which would imply a finite VEV for ¯ψγµψin the theory of Eq. (3.29).
We argued in the previous section that gmust be large for the theory described by Eq.
(3.29) to be stable. This too seems natural in light of Eq. (3. 30), because a large gmakes
51
9→
'
Figure 3.8: Radiative corrections make the effective potential Γ[ A] stable and give Aµa non-zero
VEV.
theA2term small, so that the instability created by it may be easil y controlled by the
interaction with the fermions, yielding a VEV for Aµthat lies within the energy range of
the effective theory. Figure 3.8 schematically represents h ow the radiative corrections to
the effective action give a finite VEV for Aµ.
Armed with Eq. (3.30) it would seem possible to carry out the p rogram proposed by
Bjorken, and by Kraus and Tomboulis, in order to arrive at an a pproximation of QED in
which the photons are composite Goldstone bosons. It is conc eivable that a complicated
theory of self-interacting fermions, perhaps one with non- standard kinetic terms, might sim-
ilarly yield a VEV for ¯ψi
2(γµ→
∂ν−γµ←
∂ν)ψ, allowing the project of dynamically generating
linearized gravity to go forward.
It would have been more encouraging if we had been able to obta in a non-zero/angbracketleftbig¯ψγµψ/angbracketrightbig
through a more natural mechanism than invoking an imaginary charge. Non-abelian gauge
theories (such as QCD) exhibit attraction between particle s of the same fermion number
(and, like abelian theories with imaginary charge, they exh ibit anti-screening). So far,
however, attempts to find a non-abelian gauge theory with non -zero/angbracketleftbig¯ψγµψ/angbracketrightbig
have failed,
possibly because in such theories the attraction between fe rmion and antifermion is stronger
than the attraction between fermions (see, for instance, [4 3]).
52
Chapter 4
Phenomenology of spontaneous
Lorentz violation
What lies behind the Principle of Relativity? This is a philo sophical
question, not a scientific one. You will have your own opinion ; here is ours.
We think the Principle of Relativity as used in special relat ivity rests on
one word: emptiness. Space is empty.
— Edwin F. Taylor and John A. Wheeler, Spacetime Physics , Chap. 3
4.1 Introduction
Lorentz invariance (LI), the fundamental symmetry of Einst ein’s special relativity, states
that physical results should not change after an experiment has been boosted or rotated.
In recent years, and particularly since the publication of w ork on the possibility of sponta-
neously breaking LI in bosonic string field theory ([44]), th ere has been considerable interest
in the prospect of violating LI. More recent motivations for work on Lorentz non-invariance
have ranged from the explicit breaking of LI in the non-commu tative geometries that some
have proposed as descriptions of physical space-time (see [ 45] and references therein), and
in certain supersymmetric theories considered by the strin g community ([46, 47]), to the
possibility of explaining puzzling cosmic ray measurement s by invoking small departures
from LI ([48]) or modifications to special relativity itself ([49, 50, 51]). It has also been
suggested that anomalies in certain chiral gauge theories m ay be traded for violations of
LI and CPT ([52]). Extensions of the standard model have been proposed that are meant
to capture the low-energy effects of whatever new high-energ y physics (string theory, non-
commutative geometry, loop quantum gravity, etc.) might be introducing violations of LI
53
([53]).
Our own investigation of composite massless mediators in Ch apters 2 and 3 led us to
consider the question of how a reasonable QFT might spontane ously break LI through a
timelike Lorentz vector VEV ∝an}b∇acketle{t¯ψγµψ∝an}b∇acket∇i}ht ∝ne}ationslash= 0. This breaking of LI can be thought of conceptu-
ally as the introduction of a preferred frame: the rest frame of the fermion number density.
If some kind of gauge coupling were added to the theory withou t destroying this LI breaking,
the fermion number density would also be a charge density, an d the preferred frame would
be the rest frame of a charged background in which all process es are taking place. This
allows us to make some very general remarks in Section 4.2 on t he resulting LI-violating phe-
nomenology for electrodynamics and on experimental limits to our non-Lorentz-invariant
VEV. This discussion will be based on work previously publis hed in [24].
Experimental data put very tight constraints on Lorentz vio lating operators that involve
Standard Model particles [66], but the bounds are more model -independent on Lorentz vio-
lation that appears only in couplings to gravity [67, 68]. On e broad class of Lorentz-breaking
gravitational theories are the so-called vector-tensor th eories in which the space-time met-
ricgµνis coupled to a vector field Sµthat does not vanish in the vacuum. Consideration
of such theories dates back to [69] and their potentially obs ervable consequences are ex-
tensively discussed in [70]. These theories have an unconst rained vector field coupled to
gravity. Theories with a unit constraint on the vector field w ere proposed as a means of
alleviating the difficulties that plagued the original uncon strained theories ([71]).
The phenomenology of these theories with the unit constrain t has been recently explored.
It has been proposed as a toy model for modifying dispersion r elations at high energy ([72]).
The spectrum of long-wavelength excitations is discussed i n [73], where it was found that
all polarizations have a relativistic dispersion relation , but travel with different velocities.
Applications of these theories to cosmology have been consi dered in [74, 75]. Constraints
on these theories are weak, as for instance, there are no corr ections to the Post-Newtonian
parameters γandβ([76]). The status of this class of theories, also known as “æ ther-
theories,” is reviewed in [77].
In Section 4.4 we will show that the general low-energy effect ive action at the two-
derivative level of the Goldstones of spontaneous Lorentz v iolation by a timelike vector
VEV minimally coupled to gravity corresponds to the vector- tensor theory of gravity with
the unit constraint. This will allow us to place observation al constraints of very general
54
validity on this kind of Lorentz violation, from solar syste m tests of gravity. This discussion
will be based on work previously published in [54]. Finally, in Section 4.5 we shall discuss
the physical meaning of this kind of Lorentz violation and it s relation to some other models
that have appeared recently in the literature.
4.2 Phenomenology of Lorentz violation by a background
source
Following up on the idea presented in Chapter 3, imagine that the fermions of the universe
have some interaction that plays the role of Eq. (3.17) in giv ing a VEV to ¯ψγµψ, and that
in addition they have a U(1) gauge coupling (at this stage we have abandoned the proje ct
of producing composite photons). Then the U(1) gauge field may interact with a charged
background and we would be breaking LI in electrodynamics by introducing a preferred
frame: the rest frame of the background source.
The possibility of a vacuum that breaks LI and has non-trivia l optical properties has
already been investigated in [55, 56]. This work, however, d eals with significantly more
complicated models, both in terms of the interactions that s pontaneously break LI and of
the optical properties of the resulting vacuum. To obtain a p henomenology for our own
simpler proposal, we consider a free photon Lagrangian of th e form
Lphoton
0 =−1
4F2
µν−jµAµ, (4.1)
wherejµ=e∝an}b∇acketle{t¯ψγµψ∝an}b∇acket∇i}ht, thought of as an external source. The corresponding propag ator for
the free photon is
∝an}b∇acketle{tT{Aµ(x)Aν(y)}∝an}b∇acket∇i}ht=Dµν
F(x−y) +∝an}b∇acketle{tAµ(x)∝an}b∇acket∇i}htj∝an}b∇acketle{tAν(y)∝an}b∇acket∇i}htj, (4.2)
whereDµν(x−y) is the connected photon propagator and ∝an}b∇acketle{tAµ(x)∝an}b∇acket∇i}htjis the expectation value
ofAµin the presence of the external source.
If we takejµconstant and naively attempt to calculate the classical exp ectation value of
Aµin the presence of a constant source by integrating the Green function for electrodynam-
ics, we will get a volume divergence. We may attempt to regula te this volume divergence
55
by introducing a photon mass µ, which gives the result
∝an}b∇acketle{tAµ(x)∝an}b∇acket∇i}htj=jµ
µ2. (4.3)
(It is trivial to check that this is a solution to /squareAµ−µ2Aµ=−jµ, the wave equation for
the massive photon field with a source.) This is not satisfact ory because the disconnected
term in Eq. (4.2) will be proportional to µ−4and Feynman diagrams computed with our
modified photon propagator would produce results that depen d strongly on what we took
for a regulator. In fact the mass is physical and analogous to the effective photon mass
first described by the London brothers in their theory of the e lectromagnetic behavior of
superconductors [28]. (Using the language of particle phys ics we may say that, in the
presence of a U(1) gauge field, the VEV ∝an}b∇acketle{t¯ψγµψ∝an}b∇acket∇i}htspontaneously breaks the gauge invariance
and gives a mass to the boson, as in the Higgs mechanism.)
Photons in a superconductor propagate through a constant el ectromagnetic source. In
a simplified picture, we may think of it as a current density se t up by the motion of charge
carriers of mass mand charge e, moving with a velocity u. The proper charge density is
ρ0. The proper velocity of the charge carriers is ηµ= (1,u)/√
1−u2. The source is then
jµ=ρ0ηµ=ρ0pµ/m, wherepµis the classical energy momentum of the charge carriers. We
may think of mandρ0as deriving from the solutions to the parameters in a self-co nsistent
equation such as we had in Eq. (3.25).
The canonical energy momentum Pµof the system is Pµ=mηµ+eAµ=mjµ/ρ0+eAµ.
As is discussed in the superconductivity literature (see, f or instance, Chap. 8 in [57]), the
superconducting state has zero canonical energy momentum, which leads to the London
equation
jµ=−eρ0
mAµ. (4.4)
With thisjµinserted into the right-hand side of /squareAµ=−jµ(the wave equation for the
photon field in the Lorenz gauge), we find that we have a solutio n to the wave equation of
a massiveAµwith no source and a mass µ2=eρ0/m:
/squareAµ−eρ0
mAµ= 0. (4.5)
56
If we solve for Aµin Eq. (4.4) and substitute this back into Eq. (4.2), we get th at
∝an}b∇acketle{tT{Aµ(x)Aν(y)}∝an}b∇acket∇i}ht=Dµν
F(x−y) +m2
e2j2jµjν. (4.6)
Notice that if jµ(x) is not constant, then Fourier transformation of the second term in Eq.
(4.6) will not yield, in Feynman diagram vertices, the usual energy-momentum conserving
delta function. Therefore, presumed small violations of en ergy or momentum conservation
in electromagnetic processes could conceivably be paramet rized by the space-time variation
of the background source.1
With Eq. (4.6) and a rule for external massive photon legs, on e may then go ahead and
calculate the amplitude for various electromagnetic proce sses with this modified photon
propagator, and parametrize supposed observed violations of LI (see [59, 60, 61]) by jµ. If
we can make an estimate of the size of the mass mof the background charges, experimental
limits on the photon mass ( <2×10−16eV according to [62]) will provide a limit on the
VEV of ¯ψγµψ, in light of Eq. (4.4).
There are other consequences of a VEV ∝an}b∇acketle{t¯ψγµψ∝an}b∇acket∇i}ht ∝ne}ationslash= 0 on which we may speculate. Such a
background may have cosmological effects, a line of thought t hat might connect, for instance,
with [63]. Also, it is conceivable that such a VEV might have s ome relation to the problem of
baryogenesis, since it gives the background finite fermion n umber and spontaneously breaks
CPT, a violation that can ease the Sakharov condition of ther modynamical non-equilibrium
[64, 65].
4.3 Effective action for the Goldstone bosons of spontaneous
Lorentz violation
Here we begin by considering the general low-energy effectiv e action for a theory in which
Lorentz invariance is spontaneously broken by the VEV of a Lo rentz four-vector Sµ. With
an appropriate rescaling, the VEV satisfies
∝an}b∇acketle{tSµSµ∝an}b∇acket∇i}ht= 1, (4.7)
1This line of thought could connect to work on LI violation fro m variable couplings as discussed in [58].
57
since we assume the VEV of Sµis time-like. The existence of this VEV implies that there
exists a universal rest frame (which we sometimes refer to as the preferred frame) in which
Sµ=δµ
0. When the resulting low-energy effective action is minimall y coupled to gravity,
we shall see that it simply becomes the vector-tensor theory with the unit constraint.
Objects of mass M1andM2in a system moving relative to the preferred-frame can
experience a modification to Newton’s law of gravity of the fo rm ([70, 78])
UNewton =−GNM1M2
r/parenleftbigg
1−α2
2(w·r)2
r2/parenrightbigg
, (4.8)
where wis the velocity of the system under consideration, such as th e solar-system or
Milky Way galaxy, relative to the universal rest frame. The m ain purpose of this note is to
computeα2in theories where Lorentz invariance is spontaneously brok en by the VEV of a
four-vector.
The VEV of Sµspontaneously breaks Lorentz invariance. But as rotationa l invariance is
preserved in the preferred frame, only the three boost gener ators of the Lorentz symmetry
are spontaneously broken. The low-energy fluctuations Sµ(x) which preserve Eq. (4.7) are
the Goldstone bosons of this breaking, i.e., those that sati sfy
Sµ(x)Sµ(x) = 1. (4.9)
In the preferred-frame the fluctuations can be parameterize d as a local Lorentz transforma-
tion
Sµ(x) = Λµ
0(x) =1/radicalbig
1−φ2
1
φ
, (4.10)
where φis as vector with components φ1,φ2, andφ3.
Under Lorentz transformations Sµ(x)→Λµ
νSν(x) and the symmetry is realized non-
linearly on the fields φi. Using this field Sµ(x) we may then couple the Goldstone bosons
to Standard Model fields. Since however, the constraints on L orentz-violating operators2
involving Standard Model fields are considerable [66], we in stead focus on their couplings
to gravity, which are more model-independent because they a re always present once the
Goldstone bosons are made dynamical.
2More correctly, operators that appear to be Lorentz violati ng when the Goldstone bosons φiare set to
zero.
58
The Goldstone bosons are made dynamical by adding in kinetic terms for them. Since
Lorentz invariance is only broken spontaneously, the actio n for the kinetic terms should
still be invariant under Lorentz transformations. The only interactions relevant at the
two-derivative level and not eliminated by the constraint E q. (4.9) are3
L=c1∂αSβ∂αSβ+ (c2+c3)∂µSµ∂νSν+c4Sµ∂µSαSν∂νSα. (4.11)
Expanding this action to quadratic order in φi, one finds that the four parameters cican
be chosen to avoid the appearance of any ghosts. In particula r, we require c1+c4<0.4
To leading order, the effective action for the Goldstone boso ns is:
L=1
2/summationdisplay
i=1,2,3/bracketleftig/parenleftbig
∂µφi/parenrightbig2−α/parenleftbig
∂iφi/parenrightbig2/bracketrightig
(4.12)
whereα≡(c2+c3)/c1. By inserting a plane wave ansatz, φi(xµ)∝exp/parenleftbig
iωx0−ikx3/parenrightbig
, we
see that we have 2 transverse waves, φ1andφ2, with speed v=ω/k= 1, and one longitu-
dinal wave, φ3, withv=√1 +α. Since we’ve broken LI, massless particles no longer need
to travel at light speed. For α>0, the longitudinal Goldstone boson is superluminal. We
shall return to the issue of superluminality in Section 4.5.
This agrees with the result, discussed in [30] and in Chapter 3, that spontaneous Lorentz
violation gives us not only two transverse Goldstone bosons (which we could identify as
emergent photons) but also an extra polarization with an unu sual dispersion relation. In
[30], where the Lorentz-breaking VEV was imagined to be spac elike, that extra polarization
was timelike. In our case it is a longitudinal polarization b ecause the VEV in Eq. (4.7) was
chosen to be timelike.
4.4 The long-range gravitational preferred-frame effect
With gravity present the situation is more subtle. One expec ts the gravitons to “eat”
the Goldstone bosons, producing a more complicated spectru m [79, 80]. The covariant
3The other possible term, ǫµνρσ∂µSν∂ρSσ, is a total derivative.
4Notice that in our convention Sµis dimensionless and the ci’s have mass dimension two.
59
generalization of the constraint equation becomes
gµν(x)Sµ(x)Sν(x) = 1 (4.13)
and in the action for Sµwe replace∂µ→ ∇ µ.
Note that there is no Higgs mechanism to give the graviton a ma ss. For a gauge theory
we have the covariant derivative Dµ=∂µ−ieAµ, so that (Dµφ)2gives a term proportional
φ2A2, i.e., a gauge boson mass, when ∝an}b∇acketle{tφ∝an}b∇acket∇i}ht ∝ne}ationslash= 0. For in the case of gravity coupled to a vector
field we have
∇µSν=∂µSν+ Γν
ρµSρ, (4.14)
with
Γν
ρµ=1
2/parenleftbig
∂ρhν
µ+∂µhν
ρ−∂νhρµ/parenrightbig
(4.15)
so that there is no way to get a term proportional to S2h2.
Compare this the ghost condensate mechanism described in [8 1], where L=P(X) for
X≡gµν∂µφ∂νφ. If we assume that
P′(X=c2
∗∝ne}ationslash= 0) = 0, (4.16)
then, in the preferred frame, this implies that
∝an}b∇acketle{tX∝an}b∇acket∇i}ht=c2
∗=/angbracketleftig
˙φ2/angbracketrightig
∝ne}ationslash= 0 (4.17)
and theX2term inP(X) gives a graviton mass ˙φ4h2
00. This is different from our case,
where we get five massless graviton polarizations with differ ent propagation velocities.
Going back to our model, we see that local diffeomorphisms can be used to gauge
away the three Goldstone bosons. For under a local diffeomorp hism (which preserves the
constraint Eq. (4.13)),
S′µ(x′) =∂x′µ
∂xνSν(x) (4.18)
and withx′µ=xµ+ǫµ,Sµ≡vµ+φµ,
φ′µ(x′) =φµ(x) +vρ∂ρǫµ(4.19)
60
from which we can determine ǫµto completely remove φµ. Note that in the preferred frame,
ǫican be used to remove φi. In this gauge, the constraint Eq. (4.13) reduces to
S0(x) = (1 −h00(x)/2). (4.20)
The residual gauge invariance left in ǫ0can be used to remove h00. This is an inconvenient
choice when the sources are static. In a more general frame wi th∝an}b∇acketle{tSµ∝an}b∇acket∇i}ht=vµ, obtained by a
uniform Lorentz boost from the preferred frame, the constra int Eq. (4.13) is solved by
Sµ(x) =vµ(1−vρvσhρσ(x)/2). (4.21)
Next we discuss a toy model that provides an example of a more c omplete theory, that
at low energies reduces to the theory described above with th e vector field satisfying a unit
covariant constraint (4.13).5Consider the following non-gauge-invariant theory for a ve ctor
bosonAµ,
L=−1
2gµνgρσ∇ρAµ∇σAν+λ/parenleftbig
gµνAµAν−v2/parenrightbig2. (4.22)
Fluctuations about the minimum are given by
gµν=ηµν+hµν, Aµ=vµ+ψµ. (4.23)
This theory has one massive state Φ with mass MΦ∝λ1/2v, which is
Φ =vµψµ+hµνvµvν/2. (4.24)
In the limit that λ→ ∞ this state decouples from the remaining massless states. In the
preferred frame the only massless states are hµν, andψi. Since we have decoupled the heavy
state, we should expand
A0=v+/bracketleftbig
ψ0+vh00/2/bracketrightbig
−vh00/2→v−vh00/2, (4.25)
where in the last limit we have decoupled the heavy state. Not e that this parameterization
ofA0is precisely the same parameterization that we had above for S0. In other words,
5For a related example, see [80].
61
in the limit that we decouple the only heavy state in this mode l, the field Aµsatisfies
gµνAµAν=v2, which is the same as the constraint (4.13) with Aµ→vSµ.
In the unitary gauge with φi= 0, the only massless degrees of freedom are the gravitons.
There are the two helicity modes, which in the Lorentz-invar iant limit correspond to the
two spin-2 gravitons, along with three more helicities that are the Goldstone bosons, for a
total of five. The sixth would-be helicity mode is gauged away by the remaining residual
gauge invariance.
But the model that we started from does have a ghost, since we w rote a kinetic term
forAµthat does not correspond to the conventional Maxwell kineti c action. The ghost in
the theory is A0, which in our case is massive. The presence of this ghost mean s that this
field theory model is not a good high-energy completion for th e low-energy theory involving
onlySµand gravity that we are considering in this section. We assum e that a sensible high
energy completion exists for generic values of the ci’s.
Now we proceed to compute the preferred-frame coefficient α2appearing in the modifi-
cation to Newton’s law.
The action we consider is
S=/integraldisplay
d4x√g(LEH+LV+Lgf), (4.26)
with6
LEH=−1
16πGR (4.27)
and
LV=c1∇αSβ∇αSβ+c2∇µSµ∇νSν+c3∇µSν∇νSµ+c4Sµ∇µSαSν∇νSα.(4.28)
This is the most general action involving two derivatives ac ting onSµthat contributes to the
two-point function. Note that a coefficient c3appears, since in curved space-time covariant
derivatives do not commute. Other terms involving two deriv atives acting on Sµmay be
added to the action, but they are either equivalent to a combi nation of the operators already
present (such as adding RµνSµSν), or they vanish because of the constraint Eq. (4.13). We
6The coefficients ciappearing here are related to those appearing in, for exampl e [73], by chere
i=
−cthere
i/16πG.
62
assume generic values for the coefficients cithat in the low energy effective theory give no
ghosts or gradient instabilities.
As previously discussed, Sµsatisfies the constraint (4.13). We also assume that it does
not directly couple to Standard Model fields. In the literatu re, Eq. (4.13) is enforced by
introducing a Lagrange multiplier into the action. Here we e nforce the constraint by directly
solving forSµ, as given by Eq. (4.21), and then insert that solution back in to the action to
obtain an effective action for the metric.
In our approach there is a residual gauge invariance that in t he preferred-frame corre-
sponds to reparameterizations involving ǫ0only. To completely fix the gauge we add the
gauge-fixing term
Lgf=−α
2(SρSσSµ∂µhρσ)2. (4.29)
Neglecting interaction terms, in the preferred frame the ga uge-fixing term reduces to
Lgf=−α
2(∂0h00)2. (4.30)
Physically, this corresponds in the α→ ∞ limit to removing all time dependence in h00
without removing the static part, which is the gravitationa l potential. This is a convenient
gauge in which to compute when the sources are static.
At the two-derivative level, the only effect in this gauge of t he new operators is to modify
the kinetic terms for the graviton. The dispersion relation for the five helicities will be of
the formE=β|k|, where the velocities βare not the same for all helicities and depend on
the parameters ci([73]). This spectrum is different than that which is found in the “ghost
condensate” theory, where in addition to the two massless gr aviton helicities, there exists a
massless scalar degree of freedom with a non-relativistic d ispersion relation E∝ |k|2([81]).
There exists a range for the ci’s in which the theory has no ghosts and no gradient
instabilities ([73]). In particular, for small ci’s, no gradient instabilities appear if
c1+c2+c3
c1+c4>0 andc1
c1+c4>0. (4.31)
The condition for having no ghosts is simply c1+c4<0.
The correction to Newton’s law in Eq. (4.8) is linear order in the source. Thus to
determine its size we only need to find the graviton propagato r, since the non-linearity of
63
gravity contributes at higher order in the source. In order t o compute that term we have
to specify a coordinate system, of which there are two natura l choices. In the universal
rest frame, the sources, such as the solar system or Milky Way galaxy, will be moving and
the computation is difficult. We instead choose to compute in t he rest frame of the source,
which is moving at a speed |w| ≪1 relative to the universal rest frame. Observers in
that frame will observe the Lorentz breaking VEV vµ≃(1,−w). In the rest frame of the
source, a modified gravitational potential will be generate d. Technically this is because
terms in the graviton propagator v·k≃w·kare non-vanishing. It is natural to assume
that dynamical effects align the universal rest frame where vµ=δµ
0with the rest frame of
the cosmic microwave background.
In a general coordinate system moving at a constant speed wit h respect to the universal
frame the Lorentz-breaking VEV will be a general time-like v ectorvµ. Thus we need
to determine the graviton propagator for a general time-lik e constant vµ. Since Lorentz
invariance is spontaneously broken, the numerator of the gr aviton propagator is the most
general tensor constructed out of the vectors vµ,kνand the tensor ηρσ. There are 14 such
tensors. Writing the action for the gravitons as
S=1
2/integraldisplay
d4k˜hαβ(−k)Kαβ|σρ(k)˜hσρ(k) (4.32)
it is a straightforward exercise to determine the graviton p ropagator Pby solving
Kαβ|µν(k)Pµν|ρσ(k) =1
2/parenleftig
ηρ
αησ
β+ησ
αηρ
β/parenrightig
. (4.33)
The above set of conditions leads to 21 linear equations that determine the 14 coefficients
of the graviton propagator in terms of the coefficients ciand the VEV vµ. Seven equations
are redundant and provide a non-trivial consistency check o n our calculation.
Although it is necessary to compute all 14 coefficients in orde r to invert the propagator,
here we present only those that modify Newton’s law as descri bed previously (assuming
stress-tensors are conserved for sources). These are
Pαβ|ρσ
Newton=/braceleftig
Aηαβηρσ+ B(ηαρηβσ+ηασηβρ) + C(vαvβηρσ+vρvσηαβ)
+Dvαvβvρvσ+ E(vαvρηβσ+vαvσηβρ+vβvρηασ+vβvσηαρ)/bracerightig
.(4.34)
64
We find that each of these coefficients is independent of the gau ge parameter α. We also
numerically checked that without the presence of the gauge- fixing term the propagator could
not be inverted.
To compute the preferred-frame effect coefficient α2, we only need to focus on terms in
the momentum-space propagator proportional to ( v·k)2. To leading non-trivial order in
G(v·k)2and in theci’s we obtain, from the linear combination A+ 2B+ 2C+D+ 4E,
g00= 1 + 8πGN/integraldisplayd4k
(2π)41
k2/braceleftbigg
1−8πGN(v·k)2
k21
c1(c1+c2+c3)/bracketleftbig
2c3
1+ 4c2
3(c2+c3)+
+c2
1(3c2+ 5c3+ 3c4) +c1((6c3−c4)(c3+c4) +c2(6c3+c4))/bracketrightbig/bracerightbigg
˜T00(k),(4.35)
where in the first line kis a four-vector. Next we use vµ= (1,−w), place the source at the
origin, substitute T00=Mδ(3)(x) or˜T00(k) = 2πMδ(k0) and use
/integraldisplayd3k
(2π)3kikj
k4eik·x=1
8πr/bracketleftig
δij−xixj
r2/bracketrightig
(4.36)
to obtain
g00= 1−2GNM
r/parenleftbigg
1−(w·r)2
r28πGN
2c1(c1+c2+c3)/bracketleftbig
2c3
1+ 4c2
3(c2+c3)+
+c2
1(3c2+ 5c3+ 3c4) +c1((6c3−c4)(c3+c4) +c2(6c3+c4))/bracketrightbig/parenrightbigg
,(4.37)
where we have only written those terms that give a correction to Newton’s law proportional
to [w·r/r]2. We have also assumed that |w| ≪1 so that higher powers in w·r/rcan
be neglected. The factor of 1 /c1in the preferred-frame correction to the metric arises
because when c1→0 the “transverse” components of φihave no spatial gradient kinetic
term. Similarly, the factor of 1 /(c1+c2+c3) arises because when c1+c2+c3→0 the
“longitudinal” component of φihas no spatial gradient kinetic term. Either of these cases
causes a divergence in the static limit.7
The coefficients ciredefine Newton’s constant measured in solar system experim ents and
we find that
GN=G[1−8πG(c1+c4)]≃G
1 + 8πG(c1+c4), (4.38)
7This divergence can of course be avoided by considering high er-derivative terms in the action for the
Goldstone bosons. This would then give non-relativistic di spersion relations for these modes, E∝ |k|nfor
n >1, as was the case in [81].
65
which agrees with previous computations to linear order in t heci’s after correcting for the
differences in notation [74, 77].
The experimental bounds on deviations from Einstein gravit y in the presence of a source
are usually expressed as constraints on the metric perturba tion. Since the metric is not
gauge-invariant, these bounds are meaningful only once a ga uge is specified. In the litera-
ture, the bounds are typically quoted in harmonic gauge. Her e, the preferred-frame effect is
a particular term appearing in the solution for h00. For static sources, the gauge transfor-
mation needed to translate the solution in our gauge to the ha rmonic gauge is itself static.
But since a static gauge transformation cannot change h00, we may read off the coefficient
of the preferred-frame effect in the gauge that we used.
By inspection
α2=8πGN
c1(c1+c2+c3)/bracketleftbig
2c3
1+ 4c2
3(c2+c3) +c2
1(3c2+ 5c3+ 3c4)
+c1((6c3−c4)(c3+c4) +c2(6c3+c4))/bracketrightbig
, (4.39)
which can be compared with the experimental bound |α2|<4×10−7given in [78]. After [54]
was published, Foster and Jacobson ([82]) carried out the fu ll computation of α2in terms of
theciparameters in the vector-tensor theory with the unit constr aint and confirmed that
Eq. (4.39) is correct to leading non-trivial order.
The experimental bound on α2is obtained by considering the torque that the effect in
Eq. (4.8) would exert on the plane of the orbit of a planet. For simplicity, let us consider a
circular planetary orbit of radius r, moving around the sun, whose velocity wwith respect
to the preferred frame we take to be aligned with the z-axis, as shown in Fig. 4.1. The
average torque over one orbit is
τ=−ˆxα2GNM1M2w2
4rsin 2θ0, (4.40)
whereθ0is the inclination between the plane of planet’s orbit and th e axis of w.
This torque would cause the planes or the planets in the solar system to precess at
different rates, unless all the orbital planes were perfectl y aligned or anti-aligned with the
axis of w. If we consider, for instance, the orbits of Earth and Mercur y, whose planes are
aligned to within a few degrees, and then consider Eq. (4.40) with
66
[\]
Z
qq
e
eUAq
AASODQHW
Figure 4.1: Diagram of a planet moving in a circular orbit of radius raround the sun (located at
the origin), whose velocity with respect to the preferred fr ame is w. The inclination between the
plane of the orbit and the axis of wisθ0(the minimum value of the polar angle θduring the planet’s
orbit).
•M1= solar mass
• |w| ≃10−3(the sun’s speed with respect to the CMB rest frame)
•sin 2θ0∼O(1)
then the fact that Mercury and the Earth have maintained thei r approximate alignment
over the age of the solar system ( ∼4.5×109years) gives us, roughly, the bound in the
literature of |α2|∼<10−7.
A considerably stronger constraint on the size of the ci’s can be derived from the fact
that a particle moving faster than one of the graviton polari zations would lose energy
through gravitational ˇCerenkov radiation. In particular, this gravitational ˇCerenkov radi-
ation would limit the flux of the highest-energy cosmic rays ( which are protons moving at
nearly the speed of light). Depending on the exact assumptio ns regarding the abundance
and distribution of cosmic ray sources, the resulting bound can range from G|ci|∼<10−15to
G|ci|∼<10−31([67]). These limits, however, apply only if the extra gravi ton polarizations
propagate subluminally. We will have more to say on this issu e in Section 4.5.
67/A1
Figure 4.2: Feynman diagram for the (negligible) modification to gravit y by the coupling of the
graviton to acoustic perturbations in the CMB.
4.5 A cosmic solid
We know that the effect considered in Section 4.4, the modifica tion of gravity by the presence
of a background Sµwith a rest frame, is present in nature, because the electrom agnetic
radiation in the CMB has a conserved Poynting 4-vector:
Pµ=1
8π/parenleftbig
E2+B2,2E×B/parenrightbig
. (4.41)
This background Pµmodifies gravity because gravitons can couple to acoustic pe rtur-
bations in it, as shown in Fig. 4.2. This effect is, however, co mpletely negligible, since
the characteristic energy scale of the CMB is TCMB∼2.7 K, which means that this effect is
suppressed by a factor of/parenleftbiggTCMB
MPl/parenrightbigg2
∼10−64. (4.42)
The question remains, however, whether there might be some o ther background that, unlike
the CMB, couples strongly to gravity (and only to gravity, so as to explain why it has not
been otherwise detected). The Goldstone bosons of spontane ous Lorentz violation would
correspond to the sound waves in this background, and the mod ification to gravity comes,
as it did in Fig. 4.2 from the mixing of the gravitons with thes e acoustic modes.
In [73], the authors find the propagation velocities of the fiv e graviton polarizations in
vector-tensor theories with the unit constraint. In our lan guage, these are the velocities of
68
the two usual gravitons plus the three acoustic modes in the L orentz-violating background:
2 transverse traceless metric vtt= 1/(1−c13)→1,
2 transverse Goldstones vtrv= (c1−c2
1/2 +c2
3/2)/(c14)(1−c13)→c1/(c14),
1 longitudinal Goldstone vlgt=c123(2−c14)/c14(1−c13)(2 +c13+ 3c2)
→c123/c14,(4.43)
whereci...k≡G(ci+...+ck) and where the limits correspond to vanishing ci’s. Since, for
generalci’s, there are two distinct sound speeds, one for the longitud inal and one for the
transverse modes, our Lorentz-violating background fulfil ls the canonical definition of a
solid.8The transverse sound speed is associated with a shear mode (a deformation which
alters the shape but not the volume of a body). Linear shear wa ves are absent in a fluid
(see, for instance, Chapters III and VI in [83]).
In Section 4.4 we emphasized the difference between our model , which we may now
refer to as the “cosmic solid” model, and the “ghost condensa te” of [81]. In [81], Lorentz
invariance is broken by a VEV for a spin-0 vector field Aµ=∂µφwith a single degree of
freedom, whereas in the cosmic solid model the Lorentz invar iance is broken by a spin-
1 vector field Aµwith three degrees of freedom. Therefore the ghost condensa te has a
single Goldstone boson, with non-relativistic dispersion relationsE∝ |k|2, and it gives the
graviton a mass when minimally coupled to it, whereas the cos mic solid has three Goldstone
bosons, with dispersion relations E∝ |k|, and it does not give the graviton a mass. It turns
out that if the ghost condensate is gauged (i.e., if the ghost condensate field φis minimally
coupled to a U(1) gauge field Aµ), then the two polarizations of the gauge field provide
the two extra degrees of freedom, and the resulting model is e quivalent to the cosmic solid
([84]). Whether the ghost condensate itself admits a high-e nergy completion is unresolved
(see [85, 86]).
It can be seen from Eq. (4.43) that the speeds of the Goldstone bosons can be made su-
perluminal without introducing ghosts or other obvious pro blems in the low-energy effective
action. As pointed out in Section 4.4, if the Goldstones are r equired to be subluminal, then
8This was brought to my attention by Juan Maldacena and Ian Low .
69
α2no longer gives the strongest constraint on the size of the ci’s because a far more strin-
gent limit applies, from the gravitational ˇCerenkov radiation of the highest-energy cosmic
rays. Superluminal Goldstone bosons would evade that const raint. Whether superluminal-
ity could result from a reasonable high-energy completion, and whether the initial value
problem in the low-energy effective action is well-posed in t he presence of superluminal
modes, remain open questions.
70
Chapter 5
Some considerations on the
cosmological constant problem
5.1 Introduction
Consider Einstein-Hilbert gravity as an effective theory, c ontaining all the terms compatible
with its symmetries:
S=/integraldisplay
d4x√−g/bracketleftbig
Lmatter(φ,gµν)−2Λ +M2
PlR+···/bracketrightbig
, (5.1)
whereMPl≡/radicalbig
1/8πGand
Tµν≡1√−gδSmatter
δgµν. (5.2)
The equation of motion for the metric gµνis
Rµν−1
2gµνR−Λgµν= 8πGT µν. (5.3)
We would naturally expect that
Λ∼M4
Pl∼/parenleftbig
1028eV/parenrightbig4. (5.4)
If we let
gµν=ηµν+1
MPlhµν, (5.5)
71/A1Figure 5.1: Feynman diagram of the coupling of a single graviton to the co smological constant Λ in
Eq. (5.1). The blob may also be thought of as a collection of va cuum-to-vacuum quantum processes.
wherehµνis the graviton field, then the Λ term in Eq. (5.1) gives
−√−g(2Λ) = −2Λ−Λ
MPlhµ
µ− O(h2). (5.6)
The first term in the right-hand side of Eq. (5.6) is an irrelev ant constant, but the second
gives a tadpole diagram for the graviton, as shown in Fig. 5.1 . By Eq. (5.4) we would
therefore expect this tadpole interaction to be of order M3
Pl.
Alternatively, one can think of this tadpole diagram, shown in Fig. 5.1, as the coupling
of a single graviton to the quantum-mechanical vacuum energ y. This corresponds to moving
the Λgµνin Eq. (5.3) from the left-hand to the right-hand side and thi nking of it as the
contribution to the matter Tµνfrom the quantum-mechanical vacuum energy. In quantum
field theory, each frequency mode of the free field is a simple h armonic oscillator. Therefore,
each mode has a zero-point energy E=ω/2. We clearly have to cut off the sum at some
scale, but the successes of quantum field theory so far sugges t the cut off scale can’t be
much smaller than ∼1 TeV.
In any case, we get a positive value of Λ (the “cosmological co nstant”) far, far in excess of
what observation allows. To see qualitatively the effect of l arge positive Λ, imagine vacuum
energy inside a piston. Its energy density, ρ, is constant. If the piston is pulled out, as
shown in Fig. 5.2, the total energy must increase: dE=ρdV. By energy conservation, we
must have supplied that energy when we pulled on the piston: dW=Fdℓ=pdV=−dE.
Therefore the piston would resist being pulled out: Pressure is negative, p=−ρ.
The Newtonian limit of GR for a test mass on the edge of a unifor m sphere of radius r0
gives an acceleration
g=4π
3G(ρ+ 3p)r0. (5.7)
Therefore, the quantum vacuum energy would anti-gravitate . A value of Λ as large as
72
)
GOYDFXXP
HQHUJ\
Figure 5.2: Consider a piston filled with vacuum energy, whose density is constant. By energy con-
servation, the piston must resist being pulled out, and ther efore the vacuum energy exerts negative
pressure.
what we would expect on effective theory or quantum mechanica l grounds would rip apart
the universe, preventing it from developing any structure. It was long presumed that some
unknown symmetry of quantum gravity would forbid the Λ term i n Eq. (5.1), thus naturally
making the cosmological constant zero. In Chapter 3 we discu ssed one such idea: that the
graviton was a Goldstone boson of spontaneous Lorentz viola tion, so that the broken Lorentz
invariance protected it from getting any potential at all.
Data on the accelerated expansion of the universe, however, has recently shown that
there is a small but non-zero anti-gravitating term.[87, 88 ] Two possible approaches to this
cosmological constant problem that will be of interest to us here are:
•to imagine that the true Λ is zero, but that the universe conta ins some other field,
coupled only to gravity, which accounts for the accelerated expansion.
•to imagine that the value of Λ varies over some landscape of po ssible universes, and
that we naturally happen to live in one where Λ is small enough that structure (and
therefore intelligent life) may form.
The first line of thought will lead us in Section 5.2 to conside r whether a cosmological scalar
field can have a pressure more negative than −ρ. In Section 5.3 we will consider how the
ghost condensate of [81] would behave if it were responsible for the accelerated expansion
of the universe. In Section 5.4 we will re-examine the second line of thought in light of the
proposal that other parameters besides Λ vary over the lands cape of possible universes.
73
5.2 Gradient instability for scalar models of the dark energ y
with w <−1
Matter whose equation of state satisfies w≡p/ρ < −1 violates a number of conditions,
including the weak energy condition, generally assumed to a pply to any reasonable model
of physics [89]. However, the observational data do not excl ude the possibility that the dark
energy has w<−1 ([90, 91]). The results reported in [92] indicate that −1.26<w< −0.83
at 95% confidence level. The possibility of w<−1 has been explored by numerous authors
(see, for example, [93]–[103]). These models often contain a field with an unusual kinetic
term, which is referred to as a phantom or ghost field. In this l etter we show that for w<−1,
single scalar field models of the dark energy generally have a wrong sign gradient kinetic
term for fluctuations about the homogeneous background. Thi s result is not dependent on
general relativistic effects and survives in the flat-space l imit. Spatial inhomogeneities of the
dark energy are tightly constrained by observations of the c osmic microwave background.
In our analysis we will assume a time-dependent but spatiall y homogenous scalar back-
ground, and show that for w <−1 spatial instabilities inevitably arise. Consider the
low-energy effective theory of a scalar field coupled to gravi ty:
S=/integraldisplay
d4x√−g/bracketleftbig
M2
PlR+P+UR+V Rµν(∂µφ)(∂νφ) +···/bracketrightbig
, (5.8)
whereP,U, andVare functions of the scalar field φand its derivatives. (Because of the
anti-symmetry of Rµνρσin its first two and also in its last two indices, no non-vanish ing
invariant can be formed from it using first derivatives of φ.) Naively we might expect that
the higher-dimensional couplings of φto the Ricci tensor would be suppressed by powers
of the Planck mass MPl, making them irrelevant for cosmology after the Planck epoc h.
However, such terms are generated by graphs such as that in Fi gure 5.3. Writing the metric
asgµν=ηµν+hµν/MPl, we see that scalar-graviton interactions in Feynman diagr ams are
suppressed by the Planck mass, but when these interactions a re reassembled into the Ricci
tensor that suppression is absent. That is, the higher-dime nsional terms in Eq. (5.8) will
appear suppressed only by powers of the characteristic ener gy scale of the scalar field, M,
which may be much smaller than MPl.
We neglect terms in the action (5.8) that involve higher powe rs of the Ricci tensor. The
74/A1
Figure 5.3: The effective couplings of two gravitons to several quanta of the scalar field. The shaded
region represents interactions involving only scalars.
terms we consider are ones that can generate contributions t o the stress-energy tensor Tµν
that are not suppressed by powers of MPl. SinceTµνis obtained by varying the action with
respect to the metric, terms with more than one power of Rµνρσyield contributions that
are themselves proportional to the Ricci tensor and therefo re vanish in the flat-space limit.
Assuming a spatially homogeneous background, only the time derivatives of φwill be
non-vanishing in Eq. (5.8). It may be shown that in the limit MPl→ ∞ , the term
Rµν(∂µφ)(∂νφ)Vcontributes a term to the stress-energy tensor, which can be reproduced
by an appropriate change in the function U. Therefore we may restrict ourselves to V= 0
and consider the most general Uin order to analyze the flat-space behavior of Eq. (5.8).
It is always possible to perform a rescaling of the metric in E q. (5.8)gµν→e2wgµν,
withw= log[1+U/M2
Pl], so that the Uterm in Eq. (5.8) disappears, being absorbed into a
redefinition of the Paction for the ghost scalar field. (See, for example, [104].) The action
resulting from this rescaling, up to terms suppressed by pow ers of 1/MPl, is then
S=/integraldisplay
d4x√−g/bracketleftbig
M2
PlR+P/bracketrightbig
. (5.9)
The most general Lorentz-invariant scalar Lagrangian with out higher-derivative terms
(which we will consider later) is
L=P(X,φ), (5.10)
whereX=gµν∂µφ∂νφ. (A potential term Vwould be the component of P(X,φ) that
is independent of X.) Henceforth, P′(X,φ) will denote differentiation with respect to X.
Since the scalar field φis minimally coupled to gravity in Eq. (5.9), the stress-ene rgy tensor
75
is
Tµν=−Lgµν+ 2P′(X,φ)∂µφ∂νφ , (5.11)
and
w=P(X,φ)
T00=P(X,φ)
−P(X,φ) + 2˙φ2P′(X,φ)=−1 +2˙φ2P′(X,φ)
T00. (5.12)
Forφto account for the dark energy, we must have T00>0. Then,w <−1 requires
thatP′(X,φ)<0. Letφ0=φ0(t) be a solution to the equations of motion, and consider
the fluctuations about this solution: φ=φ0+π(x,t). When expanded in π, the effective
Lagrangian will contain a term
L=−P′(X,φ)|∇π|2+···, (5.13)
which implies that for P′(X,φ)<0 there will exist field configurations with non-zero spatial
gradients that have lower energy than the homogeneous config uration.1There is no direct
connection between the sign of w+ 1 and that of the ˙ π2term in the effective Lagrangian.
IfP′(X,φ) is negative, a finite expectation value for the gradients ma y be obtained
by adding higher powers of ( ∇π)2to theπLagrangian, but this is problematic because it
gives rise to a spatially inhomogeneous ground state for the dark energy and would lead
to inhomogeneities far larger than the limit of 10−5imposed by observations of the cosmic
microwave background.2While a potential term such as m2φ2tends to confine the gradients
to regions of size 1 /m, in most models of the dark energy V′′(φ) must be small enough that
these regions are of cosmological size.
In thew<−1 case, it is possible, by adding higher-derivative terms to the Lagrangian,
to avoid having finite spatial gradients lower the energy of t he field. Consider, for example,
L=P(X,φ) +S(X,φ)(/squareφ)2(5.14)
in which case
T00=−L+ 2[P′(X,φ)˙φ2+S′(X,φ)˙φ2(∂2φ)2+ 2S(X,φ)¨φ(∂2φ)−∂0(˙φS∂2φ)].(5.15)
1Here we mean energy constructed from the Hamiltonian for fluc tations about the background field
configuration.
2A condensate of gradients with a preferred magnitude, deter mined by the higher-order terms that stabi-
lize Eq. (5.13), will spontaneously break the O(3) rotational symmetry down to O(2). The homotopy group
π2[O(3)/O(2)] is non-trivial, which leads to the formation of global m onopole (hedgehog) configurations.
76
Setting the spatial gradients of φto zero, we have that
˙φ2Cgrad−2∂0(S¨φ)˙φ=(w+ 1)
2T00, (5.16)
whereCgradis the coefficient of −(∇π)2in theπLagrangian. If ∂0(S¨φ)˙φ>0, then a model
may have both Cgrad>0 andw <−1. But for wsignificantly less than −1, this also
requires ¨φ2to be at least of order M2˙φ2, unlessS(X,φ) is made unnaturally large. It is
not clear how to treat these higher-derivative terms self-c onsistently beyond perturbation
theory, so the dynamics of such models cannot be analyzed in a straightforward manner.
The models we consider below have higher powers of first deriv atives, but they satisfy the
condition that ¨φ2≪(˙φM)2.
Our analysis shows that w<−1 scalar models typically require a wrong sign ( ∇π)2term
in the effective Lagrangian. Previous analyses of ghost mode ls ([89, 105]) have focused on
the problems associated with negative energy, particularl y with a kinetic term L=−(∂µφ)2
that has the wrong sign for boththe time and space derivatives. The classical equations
of motion for such models do not exhibit growing modes of non- zero spatial gradients,
although the energy of the field is unbounded from below. Mode ls withw<−1 that do not
have a wrong sign time-derivative kinetic term in the effecti ve Lagrangian can result from a
Lorentz-invariant action, as we demonstrate below. Howeve r, both Lorentz invariance and
time translation invariance are spontaneously broken by a t ime-dependent condensate.
In [81] a model with L=P(X) was proposed in which a ghost field has a time-dependent
condensate (from now on we take the Lagrangian to be a functio n of X only, and therefore
invariant under the shift φ→φ+c). We use units in which the dimensional scale Mof
the model is unity ( M∼10−3eV if the ghost comprises the dark energy). The flat-space
equation motion is
∂µ/bracketleftbig
P′(X)∂µφ/bracketrightbig
= 0. (5.17)
Homogeneous solutions of the equations of motion with ˙φ2=c2were considered in [81].
In general, the existence of a ˙φcondensate allows for exotic equations of state, including
w<−1. In what follows we let
P(X) =−1 + 2(X−1)2+ (X−1)3, (5.18)
77
which leads to w<−1 withT00>0 whenX <1.
The energy density is given by
T00=H=∂L
∂˙φ˙φ− L= 2˙φ2P′(X)− L, (5.19)
which is not necessarily minimized by a particular ghost con densateφ=ct, although it is
a solution to the flat-space equations of motion for any value ofc. This is possible because
there is a conserved charge associated with the shift symmet ry,
Q=/integraldisplay
d3x P′(X)˙φ , (5.20)
so configurations that do not extremize T00can still be stable. In fact, the Lagrangian
describing small fluctuations has the correct sign of ˙ π2ifP′(X) + 2XP′′(X)>0. This
condition is satisfied in the region X < 1 by (5.18) given above. There is then a local
instability to the formation of gradients, as required by ou r earlier results.
5.3 Time evolution of wfor ghost models of the dark energy
Ghost models of the dark energy that approach w=−1 asymptotically make potentially
interesting predictions for the time evolution of the equat ion of state for the dark energy.
In a FRW universe, the equation of motion for the ghost field is
∂µ/bracketleftbig
a3(t)P′(X)∂µφ/bracketrightbig
= 0, (5.21)
wherea(t) is the FRW scale factor. If there is a value c2
∗=˙φ2=Xsuch thatP′(c2
∗) = 0,
then Eq. (5.12) implies that w=−1 whenX=c2
∗. The model described by Eq. (5.18) has
c2
∗= 1, and if we apply Eq. (5.21) to it, we see that if we start from X=c2
iwithciclose
toc∗, then we are driven asymptotically towards X=c2
∗andw=−1.
In the model described by Eq. (5.18), we may be driven towards w=−1 either from
above or from below, depending on whether we chose to start fr omc2
i>1 or fromc2
i<1.
We have argued that w<−1 is problematic because of spatial gradient instabilities , so that
the case in which we are driven to w=−1 from above is more interesting.
78
Near the asymptotic value c∗= 1 we have
˙π=P′(c2
i)ci
2P′′(c2∗)c2∗/parenleftigai
a/parenrightig3
. (5.22)
Thus, in this regime,
w=−1−4P′′(c2
∗)c3
∗˙π
P(c2∗)=−1−2P′(c2
i)c∗ci
P(c2∗)/parenleftbigg1 +zi
1 +z/parenrightbigg3
. (5.23)
Equation (5.23) offers a prediction for the wparameter of the dark energy as a function of
the redshift z, which could be tested by cosmological data.
In summary, from Eqs. (5.12) and (5.13) we find that in single s calar field models
of the dark energy with w <−1, the kinetic term for fluctuations about the homogeneous
background has a wrong sign gradient term. On the other hand, there is no direct connection
between the sign of the ˙ π2kinetic term in the effective Hamiltonian and the sign of w+ 1.
5.4 Anthropic distribution for cosmological constant and p ri-
mordial density perturbations
The anthropic principle has been proposed as a possible solu tion to the two cosmological
constant problems: why the cosmological constant Λ is order s of magnitude smaller than
any theoretical expectation, and why it is non-zero and comp arable today to the energy
density in other forms of matter ([106, 107, 108]). This anth ropic argument, which predates
direct cosmological evidence of the dark energy, is the only theoretical prediction for a small,
non-zero Λ ([108, 109]). It is based on the observation that t he existence of life capable of
measuring Λ requires a universe with cosmological structur es such as galaxies or clusters of
stars. A universe with too large a cosmological constant eit her doesn’t develop any structure,
since perturbations that could lead to clustering have not g one non-linear before the universe
becomes dominated by Λ, or else has a very low probability of e xhibiting structure-forming
perturbations, because such perturbations would have to be so large that they would lie
in the far tail-end of the cosmic variance. The existence of t he string theory landscape,
in which causally disconnected regions can have different co smological and particle physics
properties, adds support to the notion of an anthropic rule f or selecting a vacuum.
79
How well does this principle explain the observed value of Λ i n our universe? Careful
analysis by [109] finds that 5% to 12% of universes would have a cosmological constant
smaller than our own. In everyday experience we encounter ev ents at this level of confi-
dence,3so as an explanation this is not unreasonable.
If the value of Λ is not fixed a priori , then one might expect other fundamental pa-
rameters to vary between universes as well. This is the case i f one sums over wormhole
configurations in the path integral for quantum gravity ([11 0]), as well as in the string
theory landscape ([111, 112, 113, 114]). In [114] it was emph asized that all the parameters
of the low energy theory would vary over the space of vacua (“t he landscape”). Douglas
([112]) has initiated a program to quantify the statistical properties of these vacua, with
additional contributions by others ([113]).
In [115], Aguirre stressed that life might be possible in uni verses for which some of the
cosmological parameters are orders of magnitude different f rom those of our own universe.
The point is that large changes in one parameter can be compen sated by changes in another
in such a way that life remains possible. Anthropic predicti ons for a particular parameter
value will therefore be weakened if other parameters are all owed to vary between universes.
One cosmological parameter that may significantly affect the anthropic argument is Q, the
standard deviation of the amplitude of primordial cosmolog ical density perturbations. Rees
in [116] and Tegmark and Rees in [117] have pointed out that if the anthropic argument is
applied to universes where Qis not fixed but randomly distributed, then our own universe
becomes less likely because universes with both Λ and Qlarger than our own are anthropi-
cally allowed. The purpose of the work in this section is to qu antify this expectation within
a broad class of inflationary models. Restrictions on the a prori probability distribution
forQnecessary for obtaining a successful anthropic prediction for Λ, were considered in
[118, 119].
In our analysis we let both Λ and Qvary between universes and then quantify the
anthropic likelihood of a positive cosmological constant l ess than or equal to that observed
in our own universe. We offer a class of toy inflationary models that allow us to restrict the
a priori probability distribution for Q, making only modest assumptions about the behavior
of the a priori distribution for the parameter of the inflaton potential in t he anthropically
allowed range. Cosmological and particle physics paramete rs other than Λ and Qare held
3For instance, drawing two pairs in a poker hand.
80
fixed as initial conditions at recombination. We provisiona lly adopt Tegmark and Rees’s
anthropic bound on Q: a factor of 10 above and below the value measured in our unive rse.
Even though this interval is small, we find that the likelihoo d that our universe has a typical
cosmological constant is drastically reduced. The likelih ood tends to decrease further if
larger intervals are considered.
Weinberg determined in [108] that, in order for an overdense region to go non-linear be-
fore the energy density of the universe becomes dominated by Λ, the value of the overdensity
δ≡δρ/ρmust satisfy
δ >/parenleftbigg729Λ
500¯ρ/parenrightbigg1/3
. (5.24)
In a matter-dominated universe this relation has no explici t time dependence. Here ¯ ρis
the energy density in non-relativistic matter. Perturbati ons not satisfying the bound cease
to grow once the universe becomes dominated by the cosmologi cal constant. For a fixed
amplitude of perturbations, this observation provides an u pper bound on the cosmological
constant compatible with the formation of structure. Throu ghout our analysis we assume
that at recombination Λ ≪¯ρ.
To quantify whether our universe is a typical, anthropicall y allowed universe, additional
assumptions about the distribution of cosmological parame ters and the spectrum of density
perturbations across the ensemble of universes are needed.
A given slow-roll inflationary model with reheating leads to a Friedman-Robertson-
Walker universe with a (late-time) cosmological constant Λ and a spectrum of perturbations
that is approximately scale-invariant and Gaussian with a v ariance
Q2≡ ∝an}b∇acketle{t˜δ2∝an}b∇acket∇i}htHC. (5.25)
The expectation value is computed using the ground state in t he inflationary era and per-
turbations are evaluated at horizon-crossing. The varianc e is fixed by the parameters of
the inflationary model together with some initial condition s. Typically, for single-field φ
slow-roll inflationary models,
Q2∼H4
˙φ2/vextendsingle/vextendsingle/vextendsingle/vextendsingle
HC. (5.26)
This leads to spatially separated over- or underdense regio ns with an amplitude δthat for
81
a scale-invariant spectrum are distributed (at recombinat ion) according to
N(σ,δ) =/radicalbigg
2
π1
σe−δ2/2σ2. (5.27)
(The linear relation between Qand the filtered σin Eq. (5.27) is discussed below.)
By Bayes’s theorem, the probability for an anthropically al lowed universe (i.e., the
probability that the cosmological parameters should take c ertain values, given that life
has evolved to measure them) is proportional to the product o f thea priori probability
distribution Pfor the cosmological parameters, times the probability tha t intelligent life
would evolve given that choice of parameter values. Followi ng [109], we estimate that
second factor as being proportional to the mean fraction F(σ,Λ) of matter that collapses
into galaxies. The latter is obtained in a universe with cosm ological parameters Λ and σ
by spatially averaging over all over- or underdense regions , so that ([109])
F(σ,Λ) =/integraldisplay∞
δmindδN(σ,δ)F(δ,Λ). (5.28)
The lower limit of integration is provided by the anthropic b ound of Eq. (5.24), which gives
δmin≡(729Λ/500¯ρ)1/3. The anthropic probability distribution is
P(σ,Λ) =P(Λ,σ)F(σ,Λ)dΛdσ . (5.29)
Computing the mean fraction of matter collapsed into struct ures requires a model for
the growth and collapse of inhomogeneities. The Gunn-Gott m odel ([120, 121]) describes
the growth and collapse of an overdense spherical region sur rounded by a compensating
underdense shell. The weighting function F(δ,Λ) gives the fraction of mass in the inhomo-
geneous region of density contrast δthat eventually collapses (and then forms galaxies). To
a good approximation it is given by ([109])
F(δ,Λ) =δ1
δ+δmin. (5.30)
Additional model dependence occurs in the introduction of t he parameter sgiven by the
ratio of the volume of overdense sphere to the volume of the un derdense shell surrounding
the sphere. We will set s= 1 throughout.
82
Since the anthropically allowed values for Λ are so much smal ler than any other mass
scale in particle physics, and since we assume that Λ = 0 is not a special point in the
landscape, we follow [122, 109] in using the approximation P(Λ)≃P(Λ = 0) for Λ within
the anthropically allowed window.4The requirement that the universe not recollapse before
intelligent life has had time to evolve anthropically rules out large negative Λ ([107, 124]).
We will assume that the anthropic cutoff for negative Λ is clos e enough to Λ = 0 that all
Λ<0 may be ignored in our calculations.
As an example of a concrete model for the variation in Qbetween universes, we consider
inflaton potentials of the form (see, for example, [125])
V= Λ +λφ2p, (5.31)
wherepis a positive integer.5We assume there are additional couplings that provide
an efficient reheating mechanism, but are unimportant for the evolution of φduring the
inflationary epoch. The standard deviation of the amplitude of perturbations gives
Q=A√
λφp+1
HC
M3
Pl, (5.32)
whereAis a constant, and φHCis the value of the field when the mode of wave number
kleaves the horizon. This φHChas logarithmic dependence on λandk, which we neglect.
Randomness in the initial value for φaffects only those modes that are (exponentially) well
outside our horizon. Throughout this section, we will set th e spectral index to 1 and ignore
its running. Equation (5.32) then gives λ∝Q2.
Next, suppose that the fundamental parameters of the Lagran grian are not fixed, but
vary between universes, as might be expected if one sums over wormhole configurations in
the path integral for quantum gravity ([110]) or in the strin g theory landscape ([111, 112,
114, 113]). To obtain the correct normalization for the dens ity perturbations observed in
our universe, the self-coupling must be extremely small. As the standard deviation Qwill
be allowed to vary by an order of magnitude around 10−5, for this model the self-coupling
in alternate universes will be very small as well.
4Garriga and Vilenkin point to examples of quintessence mode ls in which the approximation
P(Λ)≃P(Λ = 0) in the anthropically allowed range is not valid [123].
5Recent analysis of astronomical data disfavors the λφ4inflationary model ([126]), but for generality we
will consider an arbitrary pin Eq. (5.31).
83
We may then perform an expansion about λ= 0 for the a priori probability distribu-
tion ofλ. The smallness of λsuggests that we may keep only the leading term in that
expansion. If the a priori probability distribution extends to negative values of λ(which
are anthropically excluded due to the instability of the res ulting action for φ), we expect it
to be smooth near λ= 0, and the leading term in the power series expansion to be ze roth
order inλ(i.e., a constant). Therefore we expect a flat a priori probability distribution for
λ. The a priori probability distribution for Qis then
P(Q)∝dλ
dQ∼Q , (5.33)
where the normalization constant is determined by the range of integration in Q. Note that
this distribution favors large Q. On the other hand, if the a priori probability distribution
for the coupling λonly has support for λ>0 thenλ= 0 is a special point and we cannot
argue that P(Q)∝Q. However, since the anthropically allowed values of λare very small,
thea priori distribution for λshould be dominated, in the anthropically allowed window,
by a leading term such as P(λ)∼λq. Normalizability requires q>−1. Usingλ∝Q2, this
givesP(Q)∼Q2q+1.
Before proceeding, it is convenient to transform to the new v ariables:
y≡Λ
ρ∗; ˆσ≡σ/parenleftbigg¯ρ
ρ∗/parenrightbigg1
3
. (5.34)
Here ¯ρis the energy density in non-relativistic matter at recombi nation, which we take
to be fixed in all universes, and ρ∗is the value for the present-day energy density of
non-relativistic matter in our own universe. For a matter-d ominated universe ˆ σis time-
independent, whereas yis constant for any era. Here and throughout this section, a s ubscript
∗denotes the value that is observed in our universe for the cor responding quantity. The
only quantities whose variation from universe to universe w e will consider are yand ˆσ.
In terms of these variables and following [109], the probabi lity distribution of Eq. (5.29)
is found to be
P=NdˆσdyP (ˆσ)/integraldisplay∞
βdxe−x
β1/2+x1/2, (5.35)
where
β≡1
2ˆσ2/parenleftbigg729y
500/parenrightbigg2/3
, (5.36)
84
andNis the normalization constant.
Notice that, since x≥β, largeβimplies that P ∼e−β≪1. For a fixed ˆ σ, largey
implies large β. Thus, for fixed ˆ σ, large cosmological constants are anthropically disfavor ed.
But if ˆσis allowed to increase, then β∼ O(1) may be maintained at larger y. Garriga and
Vilenkin have pointed out that the distribution in Eq. (5.35 ) may be rewritten using the
change of variables (ˆ σ,y)∝ma√sto→(ˆσ,β) ([119]). The Jacobian for that transformation is a
function only of ˆ σ. Equation (5.35) then factorizes into two parts: one depend ing only on
ˆσ, the other only on β. Integration over ˆ σproduces an overall multiplicative factor that
cancels out after normalization, so that any choice of P(ˆσ) will give the same distribution
for the dimensionless parameter β. In that sense, even in a scenario where ˆ σis randomly
distributed, the computation in [109] may be seen as an anthr opic prediction for β.6The
measured value of βis, indeed, typical of anthropically allowed universes, bu t an anthropic
explanation for βalone does not address the problem of why both Λ and Qshould be so
small in our universe.
Implementing the anthropic principle requires making an as sumption about the mini-
mum mass of “stuff” collapsed into stars, galaxies, or cluste rs of galaxies that is needed for
the formation of life. It is more convenient to express the mi nimum mass Mminin terms of
a comoving scale R:Mmin= 4π¯ρa3
eqR3/3 (by convention a= 1 today, so Ris a physical
scale). We do not know the precise value of R. A better understanding of biology would
in principle determine its value, which should only depend o n chemistry, the fraction of
matter in the form of baryons, and Newton’s constant. In our a nalysis these are all fixed
initial conditions at recombination. In particular, we wou ld not expect Mminto depend on
Λ orQ.7Therefore, even though the relation between MminandRdepends on present-day
cosmological parameters, the value of this threshold will b e constant between universes be-
cause it depends only on parameters that we are treating as fix ed initial conditions. Thus,
in computing the probability distribution over universes, we will fixR. Since we don’t know
what is the correct anthropic value for R, we will present our results for both R=1 and 2
Mpc. (Ron the order of a few Mpc corresponds to requiring that struct ures as large as our
galaxy be necessary for life.)
We then proceed to filter out perturbations with wavelength s maller than R, leading to
6We thank Garriga and Vilenkin for explaining this point to us .
7Note, however, that requiring life to last for billions of ye ars (long enough for it to develop intelligence
and the ability to do astronomy) might place bounds on Q. See [117].
85
a varianceσ2that depends on the filtering scale. Expressed in terms of the power spectrum
evaluated at recombination,
σ2=1
2π2/integraldisplay∞
0dkk2P(k)W2(kR) (5.37)
whereWis the filter function, which we take to be a Gaussian W(x) =e−x2/2.P(k) is
the power spectrum, which we assume to be scale-invariant. ( ForP(k) we use Eq. (39) of
[109], setting n= 1).
Evaluating (5.37) at recombination gives, for our universe ,
ˆσ∗=C∗Q∗. (5.38)
The number C∗contains the growth factor and transfer function evaluated from horizon
crossing to recombination and only depends on physics from t hat era. We assume Λ is
small enough so that at recombination it can be ignored and th us we take the variation in
ˆσbetween universes to come solely from its explicit dependen ce onQ.
We may then use observations of Q∗andσ∗to determine ˆ σ=C∗Q, valid for all universes.
We use the explicit expression for C∗that is obtained from Eqs. (39)-(43) and (48)-(51)
in [109]. This takes as inputs the Hubble parameter H0≡100h∗km/s, the energy density
in non-relativistic matter Ω ∗, the cosmological constant λ∗= 1−Ω∗, the baryon fraction
Ωb= 0.023h−2
∗, the smoothing scale R, and the COBE-normalized amplitude of fluctuations
at horizon crossing, Q∗= 1.94×10−5Ω−.785−0.05∗ln Ω∗∗ .
As we have argued, the dependence of C∗on the cosmological constant is not relevant
for our purposes. For our calculations we use Ω ∗= 0.134h−2
∗, andh∗= 0.73, consistent
with their observed best-fit values ([127]). The smoothing s caleRwill be taken to be either
1 Mpc or 2 Mpc, and the corresponding values for C∗are 5.2·104and 3.8·104.
The values chosen for the range of Qare motivated by the discussion in [117] about
anthropic limits on the amplitude of the primordial density perturbations. The authors of
[117] argue that Qbetween 10−3and 10−1leads to the formation of numerous supermassive
blackholes, which might obstruct the emergence of life.8They then claim that universes
withQless than 10−6are less likely to form stars, or if star clusters do form, tha t they would
8They also note that for Q >10−4formation of life is possible, but planetary disruptions ca used by flybys
may make it unlikely for planetary life to last billions of ye ars.
86
P(Q)∝1/Q0.9in the range P(Q)∝Qin the range
Q∗/10<Q< 10Q∗Q∗/15<Q< 15Q∗Q∗/10<Q< 10Q∗Q∗/15<Q< 15Q∗
R= 1 Mpc R= 2 Mpc R= 1 Mpc R= 2 Mpc R= 1 Mpc R= 2 Mpc R= 1 Mpc R= 2 Mpc
P(y<y ∗) 1·10−33·10−34·10−41·10−35·10−41·10−31·10−44·10−4
∝an}b∇acketle{ty∝an}b∇acket∇i}ht/y∗ 1·1044·1034·1041·1041·1045·1034·1042·104
y5%/y∗ 9·10 4·10 3·1021·1022·1027·10 6·1022·102
Table 5.1: Anthropically determined properties of the cosmological c onstant.
not be bound strongly enough to retain supernova remnants. S ince there is considerable
uncertainty in these limits, we carry out calculations usin g both the range indicated by
[117] as well as a range that is somewhat broader.9
Previous work on applying the anthropic principle to variab le Λ andQhas assumed a
priori distributions P(Q) that fall off as 1 /Qkfor largeQ, withk≥3 [118, 119]. Such
distributions were chosen in order to keep the anthropic pro bability P(y,Q) normalizable,
and they usually yield anthropic predictions for the cosmol ogical constant similar to those
that were obtained in [109] by fixing Qto its observed value, because they naturally favor
aQas small as its observed value in our universe. For instance, forP(Q)∝1/Q3in the
rangeQ∗/10< Q < 10Q∗,P(y < y∗) = 5% for R= 1 Mpc, while P(y < y∗) = 7% for
R= 2 Mpc.)
However, if we accept the argument of Tegmark and Rees in [117 ] that there are natural
anthropic cutoffs on Q, it follows that the behavior of P(Q) at largeQis irrelevant to the
normalizability of P(y,Q). Furthermore, P(Q)∼1/Qkin the neighborhood of Q= 0 for
k≥1 leads to an unnormalizable distribution, since the integr al/integraltext
P(Q)dQblows up. In
what follows we shall consider two a priori distributions: P(Q)∝Q, andP(Q)∝1/Q0.9
inside the anthropic window, motivated by the inflationary m odels we have discussed.
The results are summarized in Table 5.1, where P(y < y∗) is the anthropic probabil-
ity that the value ybe no greater than what is observed in our own universe, ∝an}b∇acketle{ty∝an}b∇acket∇i}htis the
anthropically weighed mean value of y, andy5%is the value of ysuch that the anthropic
probability of obtaining a value no greater than that is 5%.
By comparison, for this choice of cosmological parameters, the authors of [109] find that,
forQfixed (or measured), the probability of a universe having a co smological constant no
9Notice that we are using the ranges indicated in [117] as abso lute anthropic cutoffs. Arguments like
those made in [117] introduce some correction to the approxi mation made in [109] that the probability of
life is proportional to the amount of matter that collapses i nto compact structures. Since we are largely
ignorant of what the form of this correction is, we have appro ximated it as a simple window function.
87
P(Q)∝1/Q0.9in the range P(Q)∝Qin the range
Q∗/10<Q< 10Q∗Q∗/15<Q< 15Q∗Q∗/10<Q< 10Q∗Q∗/15<Q< 15Q∗
R= 1 Mpc R= 2 Mpc R= 1 Mpc R= 2 Mpc R= 1 Mpc R= 2 Mpc R= 1 Mpc R= 2 Mpc
P(Q<Q ∗) 8·10−48·10−42·10−42·10−41·10−51·10−51·10−61·10−6
∝an}b∇acketle{tQ∝an}b∇acket∇i}ht/Q∗ 8 8 11 11 8 8 13 13
Q5%/Q∗ 4 4 6 6 5 5 8 8
Table 5.2: Anthropically determined properties of the amplitude for d ensity pertubations.
greater than our own is much higher: P(y <0.7/0.3) =.05 and 0.1, forR= 1 Mpc and
R= 2 Mpc, respectively.10
One can also ask what is the probability of observing a value f orQin the range Q∗/10<
Q<Q∗, after averaging over all possible cosmological constants . Table 5.2 summarizes the
resulting distribution in Q.
In summary, inflation and a landscape of anthropically deter mined coupling constants
provides (in some scenarios) a conceptually clean framewor k for variation between universes
in the magnitude of Q. Since increasing Qallows the probability of structure to remain
non-negligible for Λ considerably larger than in our own uni verse, anthropic solutions to
the cosmological constant problem are weakened by allowing Qas well as Λ to vary from
one universe to another.
10These numbers are taken from Table 1 in the published version of [109].
88
Chapter 6
The reverse sprinkler
Everything’s got a moral, if only you can find it.
— Lewis Carroll, Alice’s Adventures in Wonderland
This chapter is based largely on [128]. Some followups that h ave appeared since the
publication of that article include [129] and [130].
6.1 Introduction
In 1985, R. P. Feynman, one of most distinguished theoretica l physicists of his time, pub-
lished a collection of autobiographical anecdotes that att racted much attention on account
of their humor and outrageousness ([131]). While describin g his time at Princeton as a
graduate student (1939–1942), Feynman tells the following story ([132]):
There was a problem in a hydrodynamics book,1that was being discussed by all
the physics students. The problem is this: You have an S-shap ed lawn sprinkler
. ..and the water squirts out at right angles to the axis and ma kes it spin in a
certain direction. Everybody knows which way it goes around ; it backs away
from the outgoing water. Now the question is this: If you . .. p ut the sprinkler
completely under water, and sucked the water in . ..which way would it turn?
1It has not been possible to identify the book to which Feynman was referring. As we shall discuss,
the matter is treated in Ernst Mach’s Mechanik , first published in 1883 ([137]). Yet this book is not a
“hydrodynamics book” and the reverse sprinkler is presente d as an example, not a problem. In [147],
John Wheeler suggests that the problem occurred to them whil e discussing a different question in the
undergraduate mechanics course that Wheeler was teaching a nd for which Feynman was the grader.
89
Feynman went on to say that many Princeton physicists, when p resented with the
problem, judged the solution to be obvious, only to find that o thers arrived with equal
confidence at the opposite answer, or that they had changed th eir minds by the following
day. Feynman claims that after a while he finally decided what the answer should be
and proceeded to test it experimentally by using a very large water bottle, a piece of
copper tubing, a rubber hose, a cork, and the air pressure sup ply from the Princeton
cyclotron laboratory. Instead of attaching a vacuum to suck the water, he applied high air
pressure inside of the water bottle to push the water out thro ugh the sprinkler. According to
Feynman’s account, the experiment initially went well, but after he cranked up the setting
for the pressure supply, the bottle exploded, and “. .. the wh ole thing just blew glass and
water in all directions throughout the laboratory .. .” ([13 3]).
Feynman ([131]) did not inform the reader what his answer to t he reverse sprinkler
problem was or what the experiment revealed before explodin g. Over the years, and partic-
ularly after Feynman’s autobiographical recollections ap peared in print, many people have
offered their analyses, both theoretical and experimental, of this reverse sprinkler problem.2
The solutions presented often have been contradictory and t he theoretical treatments, even
when they have been correct, have introduced unnecessary co nceptual complications that
have obscured the basic physics involved.
All physicists will probably know the frustration of being c onfronted by an elementary
question to which they cannot give a ready answer in spite of a ll the time dedicated to the
study of the subject, often at a much higher level of sophisti cation than what the problem
at hand would seem to require. Our intention is to offer an elem entary treatment of this
problem, which should be accessible to a bright secondary sc hool student who has learned
basic mechanics and fluid dynamics. We believe that our answe r is about as simple as it can
be made, and we discuss it in light of published theoretical a nd experimental treatments.
2In the literature it is more usual to see this problem identifi ed as the “Feynman inverse sprinkler.”
Because the problem did not originate with Feynman and Feynm an never published an answer to the
problem, we have preferred not to attach his name to the sprin kler. Furthermore, even though it is a
pedantic point, a query of the Oxford English Dictionary suggests that “reverse” (opposite or contrary in
character, order, or succession) is a more appropriate desc ription than “inverse” (turned up-side down) for
a sprinkler that sucks water.
90
Figure 6.1: A sprinkler submerged in a tank of water as seen from above. Th e L-shaped sprinkler
is closed, and the forces and torques exerted by the water pre ssure balance each other.
6.2 Pressure difference and momentum transfer
Feynman speaks in his memoirs of “an S-shaped lawn sprinkler .” It should not be difficult,
however, to convince yourself that the problem does not depe nd on the exact shape of the
sprinkler, and for simplicity we shall refer in our argument to an L-shaped structure. In
Fig. 6.1 the sprinkler is closed: Water cannot flow into it or o ut of it. Because the water
pressure is equal on opposite sides of the sprinkler, it will not turn: there is no net torque
around the sprinkler pivot.
Let us imagine that we then remove part of the wall on the right , as pictured in Fig. 6.2,
opening the sprinkler to the flow of water. If water is flowing i n, then the pressure marked
P2must be lower than the pressure P1, because water flows from higher to lower pressure.
In both Fig. 6.1 and Fig. 6.2, the pressure P1acts on the left. But because a piece of the
sprinkler wall is missing in Fig. 6.2, the relevant pressure on the upper right part of the
open sprinkler will be P2. It would seem then that the reverse sprinkler should turn to ward
the water, because if P2is less than P1, there would be a net force to the right in the upper
part of the sprinkler, and the resulting torque would make th e sprinkler turn clockwise. If
Ais the cross section of the sprinkler intake pipe, this torqu e-inducing force is A(P1−P2).
But we have not taken into account that even though the water h itting the inside wall
of the sprinkler in Fig. 6.2 has lower pressure, it also has le ft-pointing momentum. The
91
Figure 6.2: The sprinkler is now open. If water is flowing into it, then the pressures marked P1
andP2must satisfy P1>P2.
incoming water transfers that momentum to the sprinkler as i t hits the inner wall. This
momentum transfer would tend to make the sprinkler turn coun terclockwise. One of the
reasons why the reverse sprinkler is a confusing problem is t hat there are two effects in play,
each of which, acting on its own, would make the sprinkler tur n in opposite directions. The
problem is to figure out the net result of these two effects.
How much momentum is being transferred by the incoming water to the inner sprinkler
wall in Fig. 6.2? If water is moving across a pressure gradien t, then over a differential time
dt, a given “chunk” of water will pass from an area of pressure Pto an area of pressure
P−dPas illustrated in Fig. 6.3. If the water travels down a pipe of cross-section A,
its momentum gain per unit time is AdP. Therefore, over the entire length of the pipe,
the water picks up momentum at a rate A(P1−P2), whereP1andP2are the values of
the pressure at the endpoints of the pipe. (In the language of calculus,A(P1−P2) is the
total force that the pressure gradient across the pipe exert s on the water. We obtain it by
integrating over the differential force AdP.)3
3As some readers of [128] pointed out to us ([134, 135]), this s implified discussion ignores the fact that
the cross-section of a fluid flow is not in general constant whe n a pressure gradient exists. For example, for
an ideal, incompressible fluid the velocity (and therefore, through Bernoulli’s equation, also the pressure)
must be constant inside a pipe of fixed cross-section A. In that case all of the acceleration of the fluid would
have to occur outside of the sprinkler tube, as the flow narrow s down to a cross-section A. However, if P1
is the pressure of the fluid at rest, then A(P1−P2) is still the correct expression for the rate at which the
flow is gaining momentum. In fact, the shape of the flow into the reverse sprinkler will not be relevant to
our discussion at all, as should become more clear from the di scussion of conservation of angular momentum
92
Figure 6.3: As water flows down a tube with a pressure gradient, it picks up momentum.
For steady flow, the rate A(P1−P2) must be the same rate at which the water is
transferring momentum to the sprinkler wall in Fig. 6.2, bec ause otherwise the total amount
of momentum contained in the flow of water would not be constan t. Therefore A(P1−P2)
is the force that the incoming water exerts on the inner sprin kler wall in Fig. 6.2 by virtue
of the momentum it has gained in traveling down the intake pip e.
Because the pressure difference and the momentum transfer eff ects cancel each other,
it would seem that the reverse sprinkler would not move at all . Notice, however, that we
considered the reverse sprinkler only after water was alrea dy flowing continuously into it.
In fact, the sprinkler willturn toward the water initially, because the forces will bal ance
only after water has begun to hit the inner wall of the sprinkl er, and by then the sprinkler
will have begun to turn toward the incoming water. That is, in itially only the pressure
difference effect and not the momentum transfer effect is relev ant. (As the water flow stops,
there will be a brief period during which only the momentum tr ansfer and not the pressure
difference will be acting on the sprinkler, thus producing a m omentary torque opposite to
the one that acted when the water flow was being established.)
Why can’t we similarly “prove” the patently false statement that a non-sucking sprinkler
submerged in water will not turn as water flows steadily out of it? In that case the water
is going out and hitting the upper inner wall, not the left inn er wall. It exerts a force, but
that force produces no torque around the pivot. The pressure difference, on the other hand,
conservation in Section 6.3.
93
(a) (b)
Figure 6.4: The force that pushes the water must originally come from a so lid wall. The force that
causes the water flow is shown for both the regular and the reve rse sprinklers when submerged in a
tank of water.
does exert a torque. The pressure in this case has to be higher inside the sprinkler than
outside it, so the sprinkler turns counterclockwise, as we e xpect from experience.
6.3 Conservation of angular momentum
We have argued that, if we ignore the transient effects from th e switching on and switching
off of the fluid flow, we do not expect the reverse sprinkler to tu rn at all. A pertinent
question is why, for the case of the regular sprinkler, the sp rinkler-water system clearly
exhibits no net angular momentum around the pivot (with the a ngular momentum of the
outgoing water cancelling the angular momentum of the rotat ing sprinkler), while for the
reverse sprinkler the system would appear to have a net angul ar momentum given by the
incoming water. The answer lies in the simple observation th at if the water in a tank is
flowing, then something must be pushing it. In the regular spr inkler, there is a high-pressure
zone near the sprinkler wall next to the pivot, so it is this lo wer inner wall that is doing the
original pushing, as shown in Fig. 6.4(a).
For the reverse sprinkler, the highest pressure is outside t he sprinkler, so the pushing
originally comes from the right wall of the tank in which the w hole system sits, as shown
94
Figure 6.5: A tank with an opening on its side will exhibit a flow such that t he water will have an
angular momentum with respect to the tank’s bottom, even tho ugh there is no external source of
torque corresponding to the angular momentum. The apparent paradox is resolved by noting that
the tank bottom offers no inertial point of reference, becaus e the tank is recoiling due to the motion
of the water.
in Fig. 6.4(b). The force on the regular sprinkler clearly ca uses no torque around the pivot,
while the force on the reverse sprinkler does. That the water should acquire a net angular
momentum around the sprinkler pivot in the absence of an exte rnal torque might seem a
violation of Newton’s laws, but only because we are neglecti ng the movement of the tank
itself. Consider a water tank with a hole in its side, such as t he one pictured in Fig. 6.5. The
water acquires a net angular momentum with respect to any poi nt on the tank’s bottom,
but this angular momentum violates no physical laws because the tank is not inertial: It
recoils as water flows out of it.
But there is one further complication: In the reverse sprink ler shown in Fig. 6.4, the
water that has acquired left-pointing momentum from the pus hing of the tank wall will
transfer that momentum back to the tank when it hits the inner sprinkler wall, so that once
water is flowing steadily into the reverse sprinkler, the tan k will stop experiencing a recoil
force. The situation is analogous to that of a ship inside of w hich a machine gun is fired, as
shown in Fig. 6.6. As the bullet is fired, the ship recoils, but when the bullet hits the ship
wall and becomes embedded in it, the bullet’s momentum is tra nsferred to the ship. (We
assume that the collision of the bullets with the wall is comp letely inelastic.)
If the firing rate is very low, the ship periodically acquires a velocity in a direction
opposite to that of the fired bullet, only to stop when that bul let hits the wall. Thus the
ship moves by small steps in a direction opposite that of the b ullets’ flight. As the firing
rate is increased, eventually one reaches a rate such that th e interval between successive
95
Figure 6.6: In this thought experiment, a ship floats in the ocean while a m achine gun with variable
firing rate is placed at one end. Bullets fired from the gun will travel the length of the ship and hit
the wall on the other side, where they stop.
bullets being fired is equal to the time it takes for a bullet to travel the length of the ship. If
the machine gun is set for this exact rate from the beginning, then the ship will move back
with a constant velocity from the moment that the first bullet is fired (when the ship picks
up momentum from the recoil) to the moment the last bullet hit s the wall (when the ship
comes to a stop). In between those two events the ship’s veloc ity will not change because
every firing is simultaneous to the previous bullet hitting t he ship wall.
As the firing rate is made still higher, the ship will again mov e in steps, because at the
time that a bullet is being fired, the previous bullet will not have quite made it to the ship
wall. Eventually, when the rate of firing is twice the inverse of the time it takes for a bullet
to travel the length of the ship, the motion of the ship will be such that it picks up speed
upon the first two shots, then moves uniformly until the penul timate bullet hits the wall,
whereupon the ship loses half its velocity. The ship will fina lly come to a stop when the last
bullet has hit the wall. At this point it should be clear how th e ship’s motion will change
as we continue to increase the firing rate of the gun.4
For the case of continuous flow of water in a tank (rather than a discrete flow of machine
gun bullets in a ship), there clearly will be no intermediate steps, regardless of the rate of
flow. Figure 6.7 shows a water tank connected to a shower head. Water flows (with a
4Two interesting problems for an introductory university-l evel physics course suggest themselves. One
is to show that the center of mass of the bullets-and-ship sys tem will not move in the horizontal direction
regardless of the firing rate, as one expects from momentum co nservation. Another would be to analyze this
problem in the light of Einstein’s relativity of simultanei ty.
96
Figure 6.7: A water tank is connected to a shower head, so that water flows o ut. Water in the pipe
that connects the points marked A and B has a right-pointing m omentum, but as long as that pipe
is completely filled with water there is no net horizontal for ce on the tank.
consequent linear and angular momentum) between the points marked A and B, before
exiting via the shower head. When the faucet valve is opened, the tank will experience a
recoil from the outgoing water, until the water reaches B and begins exiting through the
shower head, at which point the forces on the tank will balanc e. By then the tank will have
acquired a left-pointing momentum. It will lose that moment um as the valve is closed or
the water tank becomes empty, when there is no longer water flo wing away from A but a
flow is still impinging on B.
A. K. Schultz ([136]) argues that, at each instant, the water flowing into the reverse
sprinkler’s intake carries a constant angular momentum aro und the sprinkler pivot, and if
the sprinkler could turn without any resistance (either fro m the friction of the pivot or the
viscosity of the fluid) this angular momentum would be counte rbalanced by the angular
momentum that the sprinkler picked up as the water flow was bei ng switched on. As the
fluid flow is switched off, such an ideal sprinkler would then lo se its angular momentum and
come to a halt. At every instant, the angular momentum of the s prinkler plus the incoming
water would be zero.
Schultz’s discussion is correct: In the absence of any resis tance, the sprinkler arm itself
moves so as to cancel the momentum of the incoming water, in th e same way that the ship
in Fig. 6.6 moves to cancel the momentum of the flying bullets. Resistance, on the other
97
hand, would imply that some of that momentum is picked up not j ust by the sprinkler, but
by the tank as a whole. If we cement the pivot to prevent the spr inkler from turning at all,
then the tank will pick up all of the momentum that cancels tha t of the incoming water.
How does non-ideal fluid behavior affect this analysis? Visco sity, turbulence, and other
such phenomena all dissipate mechanical energy. Therefore , a non-ideal fluid rushing into
the reverse sprinkler would acquire less momentum with resp ect to the pivot, for a given
pressure difference, than predicted by the analysis we carri ed out in Section 6.2. Thus the
pressure-difference effect would outweigh the momentum-tra nsfer effect even in the steady
state, leading to a small torque on the sprinkler even after t he fluid has begun to hit the
inside wall of the sprinkler. Total angular momentum is cons erved because the “missing”
momentum of the incoming fluid is being transmitted to the sur rounding fluid, and finally
to the tank.
6.4 History of the reverse sprinkler problem
The literature on the subject of the reverse sprinkler is abu ndant and confusing. Ernst
Mach speaks of “reaction wheels” blowing or sucking air wher e we have spoken of regular
or reverse sprinklers respectively ([137]):
It might be supposed that sucking on the reaction wheels woul d produce the
opposite motion to that resulting from blowing. Yet this doe s not usually take
place, and the reason is obvious . .. Generally, no perceptib le rotation takes place
on the sucking in of the air . .. If . ..an elastic ball, which ha s one escape-tube, be
attached to the reaction-wheel, in the manner represented i n [Fig. 6.8(a)], and
be alternately squeezed so that the same quantity of air is by turns blown out
and sucked in, the wheel will continue to revolve rapidly in t he same direction
as it did in the case in which we blew into it. This is partly due to the fact that
the air sucked into the spokes must participate in the motion of the latter and
therefore can produce no reactional rotation, but it also re sults partly from the
difference of the motion which the air outside the tube assume s in the two cases.
In blowing, the air flows out in jets, and performs rotations. In sucking, the air
comes in from all sides, and has no distinct rotation .. .
98
(a)
(b)
Figure 6.8: Illustrations from Ernst Mach’s Mechanik ([137]): (a) Figure 153 a in the original. (b)
Figure 154 in the original. (Images in the public domain, cop ied from the English edition of 1893.)
Mach appears to base his treatment on the observation that a “ reaction wheel” is not
seen to turn when sucked on. He then sought a theoretical rati onale for this observation
without arriving at one that satisfied him. Thus the bluster a bout the explanation being
“obvious,” accompanied by the tentative language about how “generally, no perceptible
rotation takes place” and by the equivocation about how the l ack of turning is “partly due”
to the air “participating in the motion” of the wheel and part ly to the air sucked “coming
in from all sides.”
Mach goes on to say that
if we perforate the bottom of a hollow cylinder . ..and place t he cylinder on
[a pivot], after the side has been slit and bent in the manner i ndicated in
[Fig. 6.8(b)], the [cylinder] will turn in the direction of t he long arrow when
blown into and in the direction of the short arrow when sucked on. The air,
here, on entering the cylinder can continue its rotation unimpeded , and this
motion is accordingly compensated for by a rotation in the op posite direction
([138]).
This observation is correct and interesting: It shows that i f the incoming water did
not give up all its angular momentum upon hitting the inner wa ll of the reverse sprinkler,
then the device would turn toward the incoming water, as we di scussed at the beginning of
99
Section 6.3.5
In his introduction to Mach’s Mechanik , mathematician Karl Menger describes it as
“one of the great scientific achievements of the [nineteenth ] century” ([139]), but it seems
that the passage we have quoted was not well known to the twent ieth-century scientists
who commented publicly on the reverse sprinkler. Feynman ([ 131]) gave no answer to
the problem and wrote as if he expected and observed rotation . Some have pointed out,
however, that the fact that he cranked up the pressure until t he bottle exploded suggests
another explanation: that he expected rotation and didn’t s ee it. This interpretation seems
to be supported by a recent letter published by E. Creutz, who claims to have been the only
other person at the Princeton cyclotron when Feynman carrie d out his experiment ([129]).
Creutz, however, explicitly disclaims any knowledge of wha t Feynman’s own theoretical
understanding of the problem was.
In [140] and [141], the authors discuss the problem and claim that no rotation is observed,
but they pursue the matter no further. In [142], it is suggest ed that students demonstrate as
an exercise that “the direction of rotation is the same wheth er the flow is supplied through
the hub [of a submerged sprinkler] or withdrawn from the hub, ” a result that is discounted
by almost all the rest of the literature.
Shortly after Feynman’s memoirs appeared, A. T. Forrester p ublished a paper in which
he concluded that if water is sucked out of a tank by a vacuum at tached to a sprinkler then
the sprinkler will not rotate ([143]). But he also made the st range claim that Feynman’s
original experiment at the Princeton cyclotron, in which he had high air pressure in the
bottle push the water out, would actually cause the sprinkle r to rotate in the direction of
the incoming water ([143]). An exchange on the issue of conse rvation of angular momentum
between Shultz and Forrester appeared shortly thereafter ( [136, 144]). The following year
L. Hsu, a high school student, published an experimental ana lysis that found no rotation
of the reverse sprinkler and questioned (quite sensibly) Fo rrester’s claim that pushing the
water out of the bottle was not equivalent to sucking it out ([ 145]). E. R. Lindgren also
published an experimental result that supported the claim t hat the reverse sprinkler did
not turn ([146]).
5In [149], P. Hewitt proposes a physical setup identical to th e one shown in Fig. 6.8(b), and observes
that the device turns in opposite directions depending on wh ether the fluid pours out of or into it. Hewitt’s
discussion seems to ignore the important difference between such a setup and the reverse sprinkler. The
issue has recently been investigated in [130].
100
After Feynman’s death, his graduate research advisor, J. A. Wheeler, published some
reminiscences of Feynman’s Princeton days from which it wou ld appear that Feynman
observed no motion in the sprinkler before the bottle explod ed (“a little tremor as the
pressure was first applied . .. but as the flow continued there w as no reaction”) ([147]). In
1992 the journalist James Gleick published a biography of Fe ynman in which he states
that both Feynman and Wheeler “were scrupulous about never r evealing the answer to the
original question” and then claims that Feynman’s answer al l along was that the sprinkler
would not turn ([148]). The physical justification that Glei ck offers for this answer is
misleading: Gleick echoes one of Mach’s comments in [137] th at the water entering the
reverse sprinkler comes in from many directions, unlike the water leaving a regular sprinkler,
which forms a narrow jet. Although this observation is corre ct, it is not very relevant to
the question at hand.
The most detailed and pertinent work on the subject, both the oretical and experimental,
was published by Berg, Collier, and Ferrell, who claimed tha t the reverse sprinkler turns
toward the incoming water ([150, 151]). Guided by Schultz’s arguments about conservation
of angular momentum ([136]), the authors offered a somewhat c omplicated statement of the
correct observation that the sprinkler picks up a bit of angu lar momentum before reaching
a steady state of zero torque once the water is flowing steadil y into the sprinkler. When
the water stops flowing, the sprinkler comes to a halt.6
The air-sucking reverse sprinkler at the Edgerton Center at MIT shows no movement at
all ([153]). As in the setups used by Feynman and others, this sprinkler arm is not mounted
on a true pivot, but rather turns by twisting or bending a flexi ble tube. Any transient
torque will therefore cause, at most, a brief shaking of such a device. The University
of Maryland’s Physics Lecture Demonstration Facility offer s video evidence of a reverse
sprinkler, mounted on a true pivot of very low friction, turn ing slowly toward the incoming
water ([152]). According to R. E. Berg, in this particular se tup
6There are other references in the literature to the reverse s prinkler. For a rather humorous exchange, see
[154] and [155]. Already in 1990 the American Journal of Physics had received so many conflicting analyses
of the problem that the editor proposed “a moratorium on publ ications on Feynman’s sprinkler” ([156]). In
one of her 1996 columns for Parade Magazine , Marilyn vos Savant, who bills herself as having the highest
recorded IQ, offered an account of Feynman’s experiment that , she claimed, settled that the reverse sprinkler
does not move ([157]). Vos Savant’s column emphasized the co nfusion of Feynman and others when faced
with the problem, leading a reader to respond with a letter to his local newspaper in which he questioned
the credibility of physicists who address matters more comp licated than lawn sprinklers, such as the origin
of the universe ([158]).
101
while the water is flowing the nozzle rotates at a constant ang ular speed. This
would be consistent with conservation of angular momentum e xcept for one
thing: while the water is flowing into the nozzle, if you reach and stop the
nozzle rotation it should remain still after you release it. [But, in practice,] after
[the nozzle] is released it starts to rotate again” ([162]).
This behavior is consistent with non-zero dissipation of ki netic energy in the fluid flow,
as we have discussed. Angular momentum is conserved, but onl y after the motion of the
tank is taken into account.7An earlier, unpublished treatment of how dissipation cause s
a steady-state torque on the reverse sprinkler is due to Titc omb, Rueckner, and Sokol
([163]). Rueckner also reports that the behavior of a sprink ler made to suck argon gas
whose viscosity is adjusted by changing its temperature see ms to corroborate that higher
viscosity leads to a larger steady-state torque. This exper iment, however, would need to be
carried out more carefully to fully confirm this effect experi mentally ([164]).
6.5 Conclusions
We have offered an elementary theoretical treatment of the be havior of a reverse sprinkler,
and concluded that, under idealized conditions, it should e xperience no torque while fluid
flows steadily into it, but as the flow commences, it will pick u p an angular momentum
opposite to that of the incoming fluid, which it will give up as the flow ends. However, in
the presence of viscosity or turbulence, the reverse sprink ler will experience a small torque
even in steady state, which would cause it to accelerate towa rd the incoming water. This
torque is balanced by an opposite torque acting on the surrou nding fluid and finally on the
tank itself.
Throughout our discussion, our foremost concern was to emph asize physical intuition
and to make our treatment as simple as it could be made (but not simpler). A question about
what L. A. Delsasso called, according to Feynman’s recollec tion, “a freshman experiment”
([133]) deserves an answer presented in a language at the cor responding level of complication.
More important is the principle, famously put forward by Fey nman himself when discussing
7In the late 1950’s and early 1960’s, there was some interest i n the related physics problem of the so-called
putt-putt (or pop-pop) boat, a fascinating toy boat that pro pels itself by heating (usually with a candle) an
inner tank connected to a submerged double exhaust. Steam bu bbles cause water to be alternately blown
out of and sucked into the tank ([159, 160, 161]). The ship mov es forward, much like Mach described the
“reaction wheel” turning vigorously in one direction as air was alternately blown out and sucked in.
102
the spin statistics theorem, that if we can’t “reduce it to th e freshman level,” we don’t
really understand it ([165]).
We also have commented on the perplexing history of the rever se sprinkler problem, a
history that is interesting not only because physicists of t he stature of Mach, Wheeler, and
Feynman enter into it, but also because it offers a startling i llustration of the fallibility of
great scientists faced with a question about “a freshman exp eriment.”
103
Bibliography
[1] D. Griffiths, Introduction to Elementary Particle Physics , (John Wiley & Sons, 1987).
[2] M. E. Peskin and D. V. Schroeder, An Introduction to Quantum Field Theory , (Perseus
Books, 1995).
[3] S. Weinberg, The Quantum Theory of Fields , Vol. I, (Cambridge University Press, 1996).
[4] S. Weinberg and E. Witten, Phys. Lett. 96B, 59 (1980).
[5] G. ’t Hooft, Nucl. Phys. B33, 173 (1971); 35, 167 (1971).
[6] H. D. Politzer, Phys. Rev. Lett. 30, 1346 (1973).
[7] D. J. Gross and F. Wilczek, Phys. Rev. Lett. 30, 1343 (1973).
[8] L. O’Raifeartaigh and N. Straumann, Rev. Mod. Phys. 72, 1 (2000).
[9] J. D. Jackson and L. B. Okun, Rev. Mod. Phys. 73, 663 (2001) [hep-ph/0012061].
[10] R. Kraichman, MIT undergraduate thesis, 1947; Phys. Re v.98, 1118 (1955); A. Pa-
papetrou, Proc. Roy. Irish Acad. 52A, 11 (1948); S. N. Gupta, Proc. Phys. Soc. London
A65, 608 (1952); R. P. Feynman, Chapel Hill Conference, unpubli shed, 1956.
[11] S. Deser, Gen. Rel. Grav. 1, 9 (1970).
[12] R. P. Feynman, F. B. Morinigo, and W. G. Wagner, Feynman Lectures on Gravitation ,
ed. B. Hatfield, (Addison-Wesley Publishing Co., 1995).
[13] C. Becchi, A. Rouet, and R. Stora, Comm. Math. Phys. 42, 127 (1975); in Renormal-
ization Theory , eds. G. Velo and A. S. Wightman (Reidel, 1976); Ann. Phys. 98, 287
(1976); I. V. Tyutin, Lebedev Institute preprint N39 (1975) .
104
[14] P. A. M. Dirac, General Theory of Relativity (John Wiley & Sons, 1975; reprinted by
Princeton University Press, 1996).
[15] L. Rosenfeld, M´ em. Acad. Roy. Belg. 6, 30 (1930); F. Belinfante, Physica 6, 887 (1939).
[16] M. B. Green, J. H. Schwarz, and E. Witten, Superstring theory , Vol 1, (Cambridge
University Press, 1987).
[17] N. Seiberg, hep-th/0601234.
[18] A. Zee, Quantum Field Theory in a Nutshell , (Princeton University Press, 2003).
[19] E. Witten, Conference in honor of Sidney Coleman at Harv ard University, unpublished,
2005.
[20] R. B. Laughlin, Phys. Rev. Lett. 50, 1395 (1983).
[21] S. C. Zhang and J. Hu, Science 294, 823 (2001) [cond-mat/0110572].
[22] P. A. M. Dirac, Proc. Roy. Soc., 209, 291 (1951).
[23] J. D. Bjorken, Ann. Phys. (N.Y.) 24, 174 (1963).
[24] A. Jenkins, Phys. Rev. D 69, 105007, (2004) [hep-th/0311127].
[25] A. Jenkins, in Proceedings of the Third Meeting on CPT and Lorentz Symmetry , ed.
V.A. Kosteleck´ y, (World Scientific, 2005) [hep-th/040918 9].
[26] Y. Nambu, Prog. Theor. Phys. extra num. , 190 (1968).
[27] Y. Nambu, in CPT and Lorentz Symmetry II , ed. V.A. Kosteleck´ y, World Scientific,
Singapore, 2002; J. Statis. Phys. 115, 7 (2004).
[28] F. London and H. London, Proc. Roy. Soc. A149 , 71 (1935).
[29] J. D. Bjorken, hep-th/0111196.
[30] P. Kraus and E. T. Tomboulis, Phys. Rev. D 66, 045015 (2002) [hep-th/0203221].
[31] H. C. Ohanian, Phys. Rev. 184(1969) 1305;
[32] Y. Nambu and G. Jona-Lasinio, Phys. Rev. 122, 345 (1961); 124, 246 (1961).
105
[33] Y. Nambu, Phys. Rev. 117, 648 (1960).
[34] M. G. Alford, K. Rajagopal and F. Wilczek, Nucl. Phys. A638 , 515C (1998) [hep-
ph/9802284].
[35] M. G. Alford, J. A. Bowers, J. M. Cheyne and G. A. Cowan, Ph ys. Rev. D 67, 054018
(2003) [hep-ph/0210106].
[36] C. Armendariz-Picon, T. Damour and V. Mukhanov, Phys. L ett.B458 , 209 (1999)
[hep-th/9904075].
[37] R. R. Caldwell, M. Kamionkowski and N. N. Weinberg, Phys . Rev. Lett. 91, 071301
(2003) [astro-ph/0302506].
[38] S. Weinberg, The Quantum Theory of Fields , Vol. 2, (Cambridge University Press,
1996).
[39] K. I. Aoki, K. Morikawa, J. I. Sumi, H. Terao and M. Tomoyo se, Phys. Rev. D 61,
045008 (2000) [hep-th/9908043].
[40] J. Bardeen, L. Cooper, and J. Schrieffer, Phys. Rev. 106, 162 (1957); Phys. Rev. 108
1175 (1957).
[41] R. F. Streater and A. S. Wightman, PCT, Spin and Statistics, and All That , (Benjamin,
1964).
[42] F. J. Dyson, Phys. Rev. 85, 631 (1952).
[43] J. Berges and K. Rajagopal, Nucl. Phys. B 538, 215 (1999) [hep-ph/9804233].
[44] V. A. Kosteleck´ y and S. Samuel, Phys. Rev. D 39, 683 (1989);
V. A. Kosteleck´ y and R. Potting, Nucl. Phys. B359 , 545 (1991).
[45] J. Madore, gr-qc/9906059; M. R. Douglas and N. A. Nekras ov, Rev. Mod. Phys. 73,
977 (2001) [hep-th/0106048].
[46] H. Ooguri and C. Vafa, Adv. Theor. Math. Phys. 7, 53 (2003) [hep-th/0302109].
[47] A. R. Frey, JHEP 0304, 012 (2003) [hep-th/0301189].
106
[48] S. R. Coleman and S. L. Glashow, Phys. Rev. D 59, 116008 (1999) [hep-ph/9812418];
Phys. Lett. B405 , 249 (1997) [hep-ph/9703240].
[49] T. G. Pavlopoulos, Phys. Rev. 159, 1106 (1967).
[50] J. Magueijo and L. Smolin, Phys. Rev. Lett. 88, 190403 (2002) [hep-th/0112090]; Phys.
Rev. D 67, 044017 (2003) [gr-qc/0207085].
[51] G. Amelino-Camelia, Nature (London) 418, 34 (2002) [gr-qc/0207049].
[52] F. R. Klinkhamer, Nucl. Phys. B535 , 233 (1998) [hep-th/9805095];
Nucl. Phys. B578 , 277 (2000) [hep-th/9912169];
F. R. Klinkhamer and J. Nishimura, Phys. Rev. D 63, 097701 (2001) [hep-th/0006154];
F. R. Klinkhamer and C. Mayer, Nucl. Phys. B616 , 215 (2001) [hep-th/0105310];
F. R. Klinkhamer and J. Schimmel, Nucl. Phys. B639 , 241 (2002) [hep-th/0205038].
[53] D. Colladay and V. A. Kosteleck´ y, Phys. Rev. D 58, 116002 (1998) [hep-ph/9809521];
V. A. Kosteleck´ y and R. Lehnert, Phys. Rev. D 63, 065008 (2001) [hep-th/0012060].
[54] M. L. Graesser, A. Jenkins and M. B. Wise, Phys. Lett. B 613, 5 (2005) [hep-
th/0501223].
[55] A. A. Andrianov and R. Soldati, Phys. Rev. D 51, 5961 (1995); Phys. Lett. B435 , 449
(1998) [hep-ph/9804448]; A. A. Andrianov, R. Soldati and L. Sorbo, Phys. Rev. D 59,
025002 (1999) [hep-th/9806220].
[56] A. A. Andrianov, P. Giacconi and R. Soldati, JHEP 0202, 030 (2002) [hep-th/0110279].
[57] C. Kittel, Quantum Theory of Solids , (Wiley, 1963).
[58] V. A. Kosteleck´ y, R. Lehnert and M. J. Perry, astro-ph/ 0212003.
[59] S. M. Carroll, G. B. Field and R. Jackiw, Phys. Rev. D 41, 1231 (1990).
[60] V. A. Kosteleck´ y and M. Mewes, Phys. Rev. D 66, 056005 (2002) [hep-ph/0205211].
[61] O. Bertolami and C. S. Carvalho, Phys. Rev. D 61, 103002 (2000) [gr-qc/9912117];
O. Bertolami, Gen. Rel. Grav. 34, 707 (2002) [astro-ph/0012462].
107
[62] K. Hagiwara et al. [Particle Data Group Collaboration], Phys. Rev. D 66, 010001
(2002).
[63] H. M. Fried, hep-th/0310095.
[64] O. Bertolami, D. Colladay, V. A. Kosteleck´ y and R. Pott ing, Phys. Lett. B395 , 178
(1997) [hep-ph/9612437].
[65] S. M. Carroll and J. Shu, hep-ph/0510081.
[66] D. Colladay and V. A. Kosteleck´ y, Phys. Rev. D 58, 116002 (1998) [hep-ph/9809521];
V. A. Kosteleck´ y and C. D. Lane, Phys. Rev. D 60, 116010 (1999) [hep-ph/9908504];
R. Bluhm and V. A. Kosteleck´ y, Phys. Rev. Lett. 84, 1381 (2000) [hep-ph/9912542].
[67] G. D. Moore and A. E. Nelson, JHEP 0109, 023 (2001) [hep-ph/0106220].
[68] C. P. Burgess, J. Cline, E. Filotas, J. Matias and G. D. Mo ore, JHEP 0203, 043 (2002)
[hep-ph/0201082].
[69] C. M. Will and K. Nordtvedt, Jr., Astrophys. J. 177, 757 (1972); R. W. Hellings and
K. Nordtvedt, Jr., Phys. Rev. D 7, 3593 (1973).
[70] C. M. Will, Theory and Experiment in Gravitational Physics , Cambridge University
Press, Cambridge (1993).
[71] T. Jacobson and D. Mattingly, Phys. Rev. D 64, 024028 (2001); D. Mattingly and
T. Jacobson, gr-qc/0112012.
[72] T. Jacobson and D. Mattingly, Phys. Rev. D 63, 041502 (2001) [hep-th/0009052].
[73] T. Jacobson and D. Mattingly, Phys. Rev. D 70, 024003 (2004) [gr-qc/0402005].
[74] S. M. Carroll and E. A. Lim, Phys. Rev. D 70, 123525 (2004) [hep-th/0407149].
[75] E. A. Lim, Phys. Rev. D 71, 063504 (2005) [astro-ph/0407437].
[76] C. Eling and T. Jacobson, Phys. Rev. D 69, 064005 (2004) [gr-qc/0310044].
[77] C. Eling, T. Jacobson and D. Mattingly, gr-qc/0410001.
108
[78] C. M. Will, in Living Reviews in Relativity , Max Planck Institute for Gravitational
Physics, Germany (2001).
[79] R. Bluhm and A. Kosteleck´ y, hep-th/0412320; V. A. Kost eleck´ y, Phys. Rev. D 69,
105009 (2004) [hep-th/0312310].
[80] B. M. Gripaios, JHEP 0410, 069 (2004) [hep-th/0408127].
[81] N. Arkani-Hamed, H. C. Cheng, M. A. Luty and S. Mukohyama , JHEP 0405, 074
(2004) [hep-th/0312099].
[82] B. Z. Foster and T. Jacobson, Phys. Rev. D 73, 064015 (2006) [gr-qc/0509083].
[83] L. D. Landau and E. M. Lifshitz, Theory of Elasticity , 3rd ed., (Reed, 1986).
[84] H. C. Cheng, M. A. Luty, S. Mukohyama and J. Thaler, hep-t h/0603010.
[85] M. L. Graesser, I. Low and M. B. Wise, Phys. Rev. D 72, 115016 (2005) [hep-
th/0509180].
[86] D. O’Connell, hep-th/0602240.
[87] A. G. Riess et al.[Supernova Search Team Collaboration], Astron. J. 116, 1009 (1998)
[astro-ph/9805201].
[88] S. Perlmutter et al. [Supernova Cosmology Project Collaboration], Astrophys. J.517,
565 (1999) [astro-ph/9812133].
[89] S. M. Carroll, M. Hoffman and M. Trodden, Phys. Rev. D 68, 023509 (2003) [astro-
ph/0301273].
[90] S. Hannestad and E. Mortsell, Phys. Rev. D 66, 063508 (2002) [astro-ph/0205096].
[91] A. Melchiorri, L. Mersini, C. J. Odman and M. Trodden, Ph ys. Rev. D 68, 043509
(2003) [astro-ph/0211522].
[92] D. N. Spergel et al., astro-ph/0603449.
[93] R. R. Caldwell, Phys. Lett. B 545, 23 (2002) [astro-ph/9908168].
[94] V. Sahni and A. A. Starobinsky, Int. J. Mod. Phys. D 9, 373 (2000) [astro-ph/9904398].
109
[95] L. Parker and A. Raval, Phys. Rev. D 60, 063512 (1999) [gr-qc/9905031].
[96] T. Chiba, T. Okabe and M. Yamaguchi, Phys. Rev. D 62, 023511 (2000) [astro-
ph/9912463].
[97] B. Boisseau, G. Esposito-Farese, D. Polarski and A. A. S tarobinsky, Phys. Rev. Lett.
85, 2236 (2000) [gr-qc/0001066].
[98] A. E. Schulz and M. J. White, Phys. Rev. D 64, 043514 (2001) [astro-ph/0104112].
[99] V. Faraoni, Int. J. Mod. Phys. D 11, 471 (2002) [astro-ph/0110067].
[100] I. Maor, R. Brustein, J. McMahon and P. J. Steinhardt, P hys. Rev. D 65, 123003
(2002). [astro-ph/0112526].
[101] V. K. Onemli and R. P. Woodard, Class. Quant. Grav. 19, 4607 (2002) [gr-
qc/0204065].
[102] D. F. Torres, Phys. Rev. D 66, 043522 (2002) [astro-ph/0204504].
[103] P. H. Frampton, Phys. Lett. B 555, 139 (2003) [astro-ph/0209037].
[104] J. Polchinski, String Theory. Vol. 1: An Introduction To The Bosonic String ,Cam-
bridge University Press, 1998.
[105] J. Cline, S. Jeon and G. Moore, hep-ph/0311312.
[106] A.D. Linde, in The Very Early Universe , eds. G.W. Gibbons et al. (Cambridge Uni-
versity Press, Cambridge, 1983); T. Banks, Nucl. Phys. B249 , 332 (1985).
[107] J.D. Barrow and F.J. Tipler, The Anthropic Cosmological Principle (Oxford Univer-
sity Press, Oxford, 1986).
[108] S. Weinberg, Phys. Rev. Lett 59, 2067 (1987).
[109] H. Martel, P. Shapiro, and S. Weinberg, Astrophys. J. 492, 29 (1998) [astro-
ph/9701099].
[110] S. Coleman, Nucl. Phys. B307 , 867 (1988); S. Giddings and A. Strominger, Nucl.
Phys.B307 , 854 (1988).
110
[111] S. Kachru, R. Kallosh, A. Linde, and S.P. Trivedi, Phys . Rev. D 68, 0046005 (2003)
[hep-th/0301240]; L. Susskind, hep-th/0302219.
[112] M. Douglas, JHEP 0305, 046 (2003) [hep-th/0303194]; hep-th/0401004.
[113] A. Giryavets, S. Kachru, P.K. Tripathy, and S.P. Trivd edi, JHEP 0404, 003,
(2003) [hep-th/0312104]; A. Giryavets, S. Kachru, and P.K. Tripathy, hep-th/0404243;
L. Susskind, hep-th/0405189; M. Dine, E. Gorbatov, and S. Th omas, hep-th/0407043.
[114] T. Banks, M. Dine, and E. Gorbatov, hep-th/0309170.
[115] A. Aguirre, Phys. Rev. D 64083508 (2001) [astro-ph/0106143].
[116] M.J. Rees, Complexity 317 (1997); in Fred Hoyle’s Universe , eds. C. Wickramasinghe
et al. (Kluwer, Dordrecht, 2003) [astro-ph/0401424].
[117] M. Tegmark and M.J. Rees, Astrophys. J. 499, 526 (1998) [astro-ph/9709058].
[118] J. Garriga, M. Livio, and A. Vilenkin, Phys. Rev. D 61, 023503 (2000) [astro-
ph/9906210].
[119] J. Garriga and A. Vilenkin, Phys. Rev. D 67, 043503 (2003) [astro-ph/0210358].
[120] P.J.E. Peebles, Astrophys. J. 147, 859 (1967).
[121] J.E. Gunn and J.R. Gott, Astrophys. J. 176, 1 (1972).
[122] A. Vilenkin, gr-qc/9512031.
[123] J. Garriga and A. Vilenkin, Phys. Rev. D 61083502 (2000) [astro-ph/9908115].
[124] S. Weinberg, in Critical Dialogues in Cosmology , ed. N. Turok (World Scientific, 1996)
[astro-ph/9610044].
[125] E.W. Kolb and M.S. Turner, The Early Universe , (Addison-Wesley, 1990).
[126] U. Seljak et al., Phys. Rev. D 71, 103515 (2005) [astro-ph/0407372].
[127] S. Eidelman et al., Phys. Lett. B 592, 1 (2004).
[128] A. Jenkins, Am. J. Phys. 72, 1276 (2004) [physics/0312087].
111
[129] E. Creutz, Am. J. Phys. 73, 198, (2005).
[130] C. Mungan, Phys. Teach. 43, L1, (2005).
[131] R. P. Feynman, Surely You’re Joking, Mr. Feynman , (Norton, 1985), pp. 63–65.
[132] R. P. Feynman, Ibid., p. 63.
[133] R. P. Feynman, Ibid., p. 65.
[134] L. Mammel, private communication, (2004).
[135] M. Jeng, private communication, (2004).
[136] A. K. Schultz, Am. J. Phys. 55, 488 (1987).
[137] E. Mach, Die Mechanik in Ihrer Entwicklung Historisch-Kritisch Dar gestellt , (Brock-
haus, 1883). In English: The Science of Mechanics: A Critical and Historical Account of
its Development (Open Court, 1960), 6th ed., pp. 388-390.
[138] E. Mach, Ibid., p. 390.
[139] E. Mach, Op. cit. , p. v.
[140] P. Kirkpatrick, Am. J. Phys. 10, 160 (1942).
[141] H. S. Belson, Am. J. Phys. 24, 413 (1956).
[142] Proceedings of the National Science Foundation Confe rence on Instruction in Fluid
Mechanics, 5–9 September 1960, Exp. 2.2, p. II–20.
[143] A. T. Forrester, Am. J. Phys. 54, 798 (1986).
[144] A. T. Forrester, Am. J. Phys. 55, 488 (1987).
[145] L. Hsu, Am. J. Phys. 56, 307 (1988).
[146] E. R. Lindgren, Am. J. Phys. 58, 352 (1990).
[147] J. A. Wheeler, Phys. Today 42(2), 24 (1989).
[148] J. Gleick, Genius: The Life and Science of Richard Feynman (Pantheon, 1992), pp.
106–108.
112
[149] P. Hewitt, Phys. Teach. 40, 390, 437 (2002).
[150] R. E. Berg and M. R. Collier, Am. J. Phys. 57, 654 (1989).
[151] R. E. Berg, M. R. Collier, and R. A. Ferrell, Am. J. Phys. 59, 349 (1991).
[152] R. E. Berg et al., University of Maryland Physics Lectu re Demonstration Facility,
<http://www.physics.umd.edu/lecdem/services/demos/d emosd3/d3-22.htm> .
[153] MIT Edgerton Center Corridor Lab: Feynman Sprinkler,
<http://web.mit.edu/Edgerton/www/FeynmanSprinkler.h tml>.
[154] M. Kuzyk, Phys. Today 42(11), 129 (1989).
[155] R. E. Berg and M. R. Collier, Phys. Today 43(7), 13 (1990).
[156] A. Mironer, Am. J. Phys. 60, 12 (1992).
[157] M. vos Savant, Parade Magazine, Oct. 6, 1996.
[158] A. de Gruyter, Houston Chronicle, Oct. 26, 1996, p. A35 .
[159] J. S. Miller, Am. J. Phys. 26, 199 (1958).
[160] R. S. Mackay, Am. J. Phys. 26, 583 (1958).
[161] I. Finnie and R. L. Curl, Am. J. Phys. 31, 289 (1963).
[162] R. E. Berg, private communication with J. M. Dlugosz an d A. Jenkins (2004).
[163] P. Titcomb, W. Rueckner, and P. E. Sokol, unpublished, (1988).
[164] W. Rueckner, private communication, (2005).
[165] R. P. Feynman, Six Easy Pieces , (Perseus, 1994), p. xxii.