Phil Lucht Math & Physics Archive
Home / Math and Physics Files / Math / Curvilinear Systems / Tensor Doc and Support / deletion pending / dead copies

tensor analysis and curvilinear coordinates

DOCX · 4.0 MB
Open DOCX file

A long self-written paper by Phil (Rimrock Digital Technology), last updated March 3, 2012, presented informally like lectures. It develops tensor analysis, metric tensor, base vectors and Standard Notation, then derives divergence, gradient, Laplacian, curl and vector Laplacian in general curvilinear coordinates. Appendices cover N-pipeds, elliptical polar coordinates, Levi-Civita tensor and dyadics. This file sits in a dead copies folder marked for deletion.

AI-written summary; may contain errors.

Extracted text (machine-read; may contain errors)
Tensor Analysis and Curvilinear Coordinates Phil Lucht Rimrock Digital Technology, Salt Lake City, Utah 84103 last update: March 3, 2012 Error! No table of contents entries found. Overview and Summary This paper develops elementary tensor analysis (also known as tensor algebra or tensor calculus) starting from Square Zero which is an arbitrary invertible continuous transformation x' = F(x) in N dimensions. The subject was "exposed" by Gregorio Ricci in the late 1800's under the name "absolute differential calculus". He and his student Tullio Levi-Civita published a masterwork on the subject in 1900 (see References). Christoffel and others had laid the groundwork a few decades earlier. The general mathematical classification of this subject is now called differential geometry. Two somewhat different applications of tensor analysis are treated concurrently. One is the subject of curvilinear coordinates in N dimensions, while the other involves transformations connecting "frames of reference". These transformations could be spatial rotations, the Lorentz transformations of special relativity, or the transformations involving the effects of gravity in general relativity. Beyond establishing the tensor analysis formalism, not much is said about this second set of applications. On the other hand, all the basic expressions for the standard differential operators in general curvilinear coordinates are derived from scratch. These results are often stated but not so often derived. The first six sections develop the theory of tensor analysis in a simple developmental notation where all indices are subscripts, just as in normal college physics. After providing motivation, the seventh section translates this developmental notation to the Standard Notation in use today. The eighth section treats transformations of length, area and volume and then the curvilinear differential operator expressions are derived, one per section, with a summary in the final section. The information is presented informally as if it were a set of lectures. Little attention is paid to mathematical rigor. There is no attempt to be concise: examples are given, tangential remarks are inserted, almost all claims are derived in line, and there is a certain amount of repetition. The material is presented in a planned sequence to minimize the need for forward references, but the sequence is not perfect. The interlocking pieces of tensor analysis do seem to exhibit a certain logical circularity. Section 1 introduces the notion of the general invertible transformation x' = F(x) as a mapping between x-space and x'-space. The range and domain of this mapping are considered in the familiar examples of polar and spherical coordinates. These same examples are used to illustrate the general ideas of coordinate lines and level surfaces. Certain Pictures are introduced to allow different names for the two inter-mapped spaces, for the function F, and for its associated objects. Section 2 introduces the linear transformations R and S=R-1 which approximate the (generally non-linear) x' = F(x) in the local neighborhood of a point x. It is shown that two types of vectors naturally arise in the context of this linearization, called contravariant and covariant, and an overbar is used to distinguish a covariant vector. Vector fields are defined and their transformations stated. The idea of scalars and vectors as tensors of rank 0 and rank 1 is presented. Section 3 defines the tangent base vectors en(x) which are tangent to the x'-coordinate lines in x-space. In the example of polar coordinates it is shown that er = and eθ = r . The vectors en exist in x-space and form there a complete basis which in general is non-orthogonal. The tangent base vectors u'n(x') of the inverse transformation x = F-1(x') are also defined. Section 4 is brief review of the notions of norm, metric and scalar product in Cartesian Space. Section 5 addresses the metric tensor, called in x-space and ' in x'-space. The metric tensor is first defined as a matrix object , and then g ≡ -1. A definition is given for two kinds of (pure) rank-2 tensors (both matrices), and it is then shown that transforms as a covariant rank-2 tensor while g is a contravariant rank-2 tensor. It is shown how applied to a contravariant vector V produces a vector that is covariant = V, and conversely g = V. In Cartesian space g = 1, so the two types of vectors coincide. The role of the metric tensor in the covariant vector dot product is stated, and the metric tensor is related to the tangent base vectors of Section 3. The Jacobian J and associated functions are defined, though the significance of J is deferred to Section 8. The last two subsections briefly discuss the connection between tensor analysis and special and general relativity with a mention of spinor algebra. Section 6 introduces the reciprocal (dual) base vectors En which are later called en in the Standard Notation. Of special interest are the covariant dot products among the en and En. It is shown how an arbitrary vector can be expanded onto different basis sets. It is found that when a contravariant vector in x-space is expanded on the en, the vector components in the expansion are in fact those of the contravariant vector in x'-space, V'i = RijVj. This fact proves useful in later sections which express differential operators in x-space in terms of curvilinear coordinates and objects of x'-space. The reciprocal base vectors U'n of the inverse transformation are also discussed. Section 7 motivates and then makes the transition from the developmental notation to the Standard Notation where contravariant indices are up and covariant ones are down. Although such a transition might seem completely trivial, confusing issues do arise. Once a matrix can have up and down indices, matrix multiplication and other matrix operations become hazy: a matrix becomes four different matrices. The matrices R and S act like tensors, but are not tensors, and in fact are not even located in a well-defined space. The third last subsection discusses the significance of tensor analysis with respect to physics in terms of covariant equations, and the second last broaches the topic of the covariant derivative of a vector field with its associated Christoffel symbols. Finally, the last subsection shows how to expand tensors of any rank in various bases and notations, and considers the expression of the non-tensor v in curvilinear coordinates. The focus then fully shifts to curvilinear coordinates as an application of tensor analysis. The final sections are all written in the Standard Notation. Section 8 shows how differential length, area and volume transform under x' = F(x). This section considers the inverse mapping of a differential orthogonal N-piped (N dimensional parallelepiped) in x'-space to a skewed one in x-space. It is shown how the scale factors h'n = describe that ratio of N-piped edges, while the Jacobian J = describes the ratio of N-piped volumes. The relationship between the vector areas of the N-pipeds is more complicated, and it is found that the ratio of vector area magnitudes is . Heavy use is made of the results of Appendices A and B, as outlined below. Sections 9 through 13 use the information of Section 8 and earlier material to derive expressions for all the standard differential operators expressed in general non-orthogonal curvilinear coordinates: divergence, gradient, Laplacian, curl, and vector Laplacian. The last two operators are treated only in N=3 dimensions where the curl has a vector representation, but then the curl is generalized to N dimensions. Section 14 summarizes all the differential operator expressions in a set of tables, and revisits the polar coordinates example one last time to illustrate a reasonably clean and practical curvilinear notation. Appendix A develops an alternative expression for the reciprocal base vector En as a generalized cross product of the tangent base vectors en, applicable when x-space is Cartesian. This alternate En is shown to match the En defined in Section 6, and the covariant dot products involving En and en are verified. Appendix B presents the geometry of a parallepiped in N dimensions (called an N-piped). Using the alternate expression for En developed in Appendix A, it is shown that the vector area of the nth pair of faces on an N-piped spanned by the en is given by ± An, where An = |det(S)| En , revealing a geometric significance of the reciprocal base vectors. Scaled by differentials so dAn = |det(S)| En(Πi≠n dx'i), this equation is then used in Section 9 where the divergence of a vector field is defined as the total flux of that field flowing out through all the faces of the skewed differential N-piped in x-space divided by its volume. This same dAn appears in Section 8 with regard to the transformation of N-piped face vector areas. Appendix C presents a case study of an N=2 non-orthogonal coordinate system, elliptical polar coordinates. Both the forward and inverse coordinate lines are displayed. The meaning of the curvilinear (x'-space) component V'n of a contravariant vector is explored in the context of this system, and the difficulties of drawing such components in non-Cartesian (curvilinear) x'-space are pondered. Finally, the Jacobian Integration Rule for changing integration variables is derived. Appendix D discusses vector densities and the Levi-Civita ε tensor, and includes a derivation of all the εε contraction formulas and their covariant statements. Appendix E describes direct product and polyadic notations (including dyadics) and shows how to expand tensors of arbitrary rank on an arbitrary basis. Appendix F then applies the methods of Appendix E to the expression of the rank-2 tensor-like object (v) in general curvilinear coordinates. The last section does the same for div(T) where T is a matrix. The objects (v) and div(T) appear in continuum mechanics. Notations diag(a,b,c..) means a diagonal matrix with diagonal elements a,b,c.. RHS, LHS refer to the right hand side and left hand side of an equation QED = which was to be demonstrated ("thus it has been proved") det(A), AT = determinant of the matrix A, transpose of a matrix A = unit vector pointing along the nth positive axis of some coordinate system // indicates a comment on something shown to the left of // Maple = a computer algebra system similar to Mathematica 1. The Transformation F: invertibility, coordinate lines, and level surfaces If x and x' are elements of the vector space RN (N-dimensional reals) , one can specify a mapping x' = F(x) F: RN → RN defined by a set of N continuous (C2) functions Fi , each of N variables, x'1 = F1(x1, x2, x3... xN) x'2 = F2(x1, x2, x3... xN) ... x'N = FN(x1, x2, x3... xN) If all functions Fi are linear in all of their arguments, then the mapping F: RN → RN is a linear mapping. Otherwise the mapping is non-linear. A mapping is often referred to as a transformation. We shall be interested only in transformations which are 1-to-1 and are therefore invertible. For such transformations, x' = F(x) x = F-1(x') , or in an equivalent notation x' = x'(x) x = x(x') In the transformation x' = F(x), if x roams over the entire RN of x-space (the domain is RN), we may find that x' roams over only some subset of RN in x'-space. The 1-to-1 invertible mapping is then between the domain of mapping F which is all of RN, and the range of mapping F which is this subset. As just noted, it will be assumed that x' = F(x) is essentially invertible so x = F-1(x') exists for any x'. By essentially is meant there may be a few problem points in the transformation which can be "fixed up" in some reasonable manner so that x' = F(x) is invertible. The functions Fi must be C1 continuous to support the linearization derivatives appearing in Section 2, and they must be C2 continuous to support some of the differential operators expressed in curvilinear coordinates in Sections 9-14 and the covariant derivative in Section 7 (v). Example 1: Polar coordinates (N=2) (a) The transformation from Cartesian to polar coordinates is given by, x = (x1, x2 ) = (x,y) x' = (x1', x2') = (θ,r) // note that r = x2' x = F-1(x') ↔ x = rcos(θ) x1 = x2' cos(x1') y = rsin(θ) x2 = x2' sin(x1') x' = F(x) ↔ r = x2' = θ = tan-1(y/x) x1' = tan-1(x2/x1) (b) The transformation is non-linear because at least one component function( e.g., r = ) is not of the form r = Ax + By. In this transformation all functions are non-linear. (c) Here is a drawing showing the nature of this mapping: The domain of x' = F(x) in x-space on the right is all of R2, but the range in x'-space is shown in gray. Imitating the language of complex variables, we can regard this gray range as depicting the principle branch of the multi-variable function x' = F(x). Other branches are obtained by shifting the gray rectangle left or right by multiples of 2π. Still other branches are obtained by taking the other branch of the real function r = which produces down-facing rectangles. The principle branch plus all the other branches then fill up the E2 of x'-space, but we care only about the principle branch range shown in gray. (d) This mapping illustrates a "problem point" involving θ = tan-1(y/x). This occurs when both x and y are 0, indicated by the red dot on the right. The inverse mapping takes the entire red line segment into this red origin point, so we have a lack of 1-to-1 going on here, meaning that formally the function F is not invertible. This can be fixed up by eliminating the red line segment from the range of F, retaining only the point at its left end. Another problem is that both the left and right vertical edges of the gray area map into the real axis in x-space, and that is fixed by removing the right edge. Thus, by doing a suitable trimming of the range, F can be made fully invertible. No one has ever had major problems using polar coordinates due to these minor issues. Example 2: Spherical coordinates (N=3) (a) The transformation from Cartesian to spherical coordinates is given by, x = (x1,x2,x3 ) = (x,y,z) x' = (x1',x2',x3') = (r,θ,φ) x = F-1(x') ↔ x = r sin(θ) cos(φ) x1 = x1'sin(x2')cos(x3') y = r sin(θ) sin(φ) x2 = x1'sin(x2')sin(x3') z = r cos(θ) x3 = x1'cos(x2') x' = F(x) ↔ r = x1' = θ = cos-1(z/) x2' = cos-1(x3/ ) φ = tan-1(y/x) x3' = tan-1(x2/x1) (b) The transformation is non-linear because at least one component function( e.g., r = ) is not of the form r = Ax + By + Cz. In this transformation, all three functions are non-linear. (c) Here is a drawing showing the nature of this mapping The domain of x' = F(x) in x-space on the right is all of E3, but the range in x'-space is the interior of an infinitely tall rectangular solid on the left we shall call an "office building". We could regard this office building as depicting the principle branch of the multi-variable function x' = F(x). Other branches are obtained by shifting the building left and right by multiples of 2π, or fore and aft by multiples of π, or by flipping it vertically, taking the other branch of r = . The principle branch plus all the other branch offices then fill up the E3 of x-space, but we care only about the principle branch office building whose walls are mostly shown in gray. (d) This mapping illustrates some "problem points". One is that entire green office building main floor (r=0) maps into the origin in x-space. This problem is fixed by trimming away the main floor keeping only the origin point of the bottom face of the office building. Another problem is that the entire red line segment (θ = 0) maps into the red point shown in x-space. This is fixed by throwing out the back wall of the office building, retaining only a line going up the left edge of the back wall. A similar problem happens on the front wall (θ = π, blue) and we fix it the same way: throw out the wall but maintain a thin line which is the left edge of this front wall (this line is missing its bottom point). Thus, by doing a suitable trimming of the range, F is made fully invertible. Cartesian Space and Quasi-Cartesian Space (a) Cartesian Space. For the purposes of this document, a Cartesian Space in N dimensions is "the usual" Hilbert Space EN in which the distance between two vectors is given by the formula d(x,y) = => [d(x+dx,x)]2 = Σi=1N (dxi)2 metric tensor = diag(1,1,1....1) as discussed in Section 4 below. The θ-r space in the above Example 1 would be a Cartesian space if it were declared that the distance between two points there was D'2 = (θ-θ')2 + (r-r')2, but that is not the usual intent in using that space. As shown below, the metric tensor used there is g = diag (r2,1) and not diag(1,1). One might argue that our Cartesian Space is in fact a Euclidean space (hence EN) having Cartesian coordinates. A non-Cartesian space is sometimes referred to as a "curved space" (non-Euclidean) and the coordinates in such a space as "curvilinear coordinates". An example is the θ-r space above. With the Cartesian Space metric tensor as gC = 1 = diag(1,1....1), the above equations can be written d2(x,y) = gCij(xi-yi)(xj-yj) and [d(x+dx,x)]2 = gCij dxi dxj ≡ (ds)2 where repeated indices are implicitly summed (sometimes called the Einstein convention). (b) Quasi-Cartesian Space. We now define a Quasi-Cartesian Space (not an official term) as one which has a diagonal metric tensor G whose diagonal elements are independently +1 or -1 instead of all +1 as with gC. In a Quasi-Cartesian Space the two equations above become d2(x,y) = Gij(xi-yi)(xj-yj) and [d(x+dx,x)]2 = Gij dxi dxj ≡ (ds)2 and of course this allows for the possibility of a negative distance squared (see Section 5 (i)). Notice that G-1 = G for any distribution of the ±1's in G. As shown later, this means that that covariant and contravariant versions of G are the same. The motivation for introducing this Quasi-Cartesian Space is to cover the case of special relativity which involves 4 dimensional linear transformations with G = diag(1,-1,-1,-1). Pictures A,B,C and D We shall always work with one of four different "pictures" involving transformations. In each picture the spaces and transformations (and their associated objects) have certain names that prove useful in certain situations. The matrices R and S are associated with transformation F as described in Section 2 below, while G and g's are metric tensors. Systems not marked Cartesian could of course be Cartesian, but we think of them as general "curved" systems with strange metric tensors. And in general, all the full transformations might be non-linear. The polar coordinates example above was presented in the context of Picture B. Picture B is the right picture for studying curvilinear coordinates where for example x-space = Cartesian coordinates and x'-space = toroidal coordinates. Picture C is useful for making statements applying to objects in curved x-space where we don't want lots of primes floating around. Pictures A and D are appropriate for consideration of general transformations, as well as linear ones like rotations and Lorentz transformations. In Sections 9-14 Picture M&S (Moon & Spencer) is introduced for the special purpose of displaying the differential operator expressions. This is Picture B with x'→ u and g'→g on the left side. The entire rest of this section uses the Picture B context. Coordinate Lines Suppose in x'-space one varies a single coordinate, say x'i, keeping all the other coordinates fixed. In x'-space the locus of points thus created is just a straight line parallel to the x'i axis, or for a principle branch situation like that of the above examples, a straight line segment. When such a straight line or segment is mapped into x-space, the result is a curve known as a coordinate line. A coordinate line is associated with a specific x'-space coordinate x'i, so one might refer to the " x'i -coordinate line", x'i being a label. In N dimensions, a point x in x-space lies on a unique set of N coordinate lines with respect to a transformation F. Remember that each such line is associated with one of the x'i coordinates. In x'-space, a point x' lies on a unique intersection of straight lines or segments, and then this all gets mapped into x-space where point x = F-1(x') then lies on a unique intersection of coordinate lines. For example, in spherical coordinates we start with some (x,y,z) in x-space and compute the xi' = (r,θ,φ) in x'-space. Our point x in x-space then lies on the r-coordinate line whose label is r, it lies on the θ-coordinate line whose label is θ, and it lies on the φ-coordinate line whose label is φ (see below). In general a coordinate "line" is some non-planar curve in N-dimensional x-space, meaning that a coordinate line might not lie on an N-1 dimensional plane. In the 2D polar coordinates example below, the red coordinate line does not lie on a 1-dimensional plane (line). In the next example of 3D spherical coordinates, it happens that every coordinate line does lie on a 2-dimensional plane. But in ellipsoidal coordinates, another 3D orthogonal system, every coordinate line does not lie on a 2-dimensional plane. Some authors refer to coordinate lines as level curves, especially in two dimensions mapping the real and imaginary part of analytic functions w = f(z) ( Ahlfors p 89). Example 1: Polar coordinates, coordinate lines Here are some coordinate lines for our prototype N=2 non-linear transformation, Cartesian to polar coordinates: The red circle is a θ-coordinate line, and the blue ray is an r-coordinate line Example 2: Spherical coordinates, coordinate lines These coordinate lines are generated exactly as described above. In x'-space one holds two coordinates fixed while allowing one to vary. The locus in x'-space is a line segment or a half line (in the case of varying r). In x-space, the corresponding coordinate lines are as shown. The green coordinate line is a θ-coordinate line, since only θ is varying. The red coordinate line is an r-coordinate line, since only r is varying. The blue coordinate line is a φ-coordinate line, since only φ is varying. The point x indicated by a black dot in x-space lies on the unique set of coordinates lines shown. Appendix C gives an example of coordinate lines for a non-orthogonal 2D coordinate system. Level Surfaces (a) Suppose in x'-space one fixes one coordinate, say x'i, and varies all the other coordinates. In x'-space the locus of points thus created is just an (N-1 dimensional) plane perpendicular to the xi axis, or for a principle branch situation like that above, a rectangle or half strip in the case of r. Mapping this planar surface in x'-space into x-space produces a surface in x-space (of dimension N-1) called a level surface. The equations of the N different xi level surface types are a'i(n) = Fi(x1, x2.....xN) i = 1,2...N where a'i(n) is some constant value selected for fixed coordinate x'i. By taking some set of closely spaced values for this constant, { a'i(1), a'i(2).....}, one obtains a family of level surfaces all of the same general shape which are closely spaced. For some different value of i, the shapes of such a family of level surfaces will in general be different. In general if f(x1, x2.....xN) = k, the set of points x which make this equation true for some fixed k is called a level set, so a level set is a surface of dimension N-1. Thus, all our level curves are also level sets. In the polar coordinates example, since there are only 2 coordinates, there is no distinction between a level surface and a coordinate line. In the spherical coordinates example, there is a distinction. If one fixes r and varies θ and φ over their horizontal rectangle inside the office building, the level surface in x-space is a sphere. If one fixes θ and varies r and φ over a left-right vertical strip inside the office building, the level surface in x-space is a sphere is a polar cone If one fixes φ and varies r and θ over a fore-aft vertical strip inside the office building, the level surface in x-space is a half plane at azimuth φ. (b) In N dimensions there will be N level surfaces in x-space, each formed by holding some x'i fixed. The intersection of N-1 level surfaces (omitting say the x3' level surface) will have all of the x'i fixed except x'3. But this describes the x'3 coordinate line. Thus, each coordinate line can be considered as the intersection of the N-1 level surfaces associated with the other coordinates. One can see this happening on the spherical coordinates example: The green coordinate line is the intersection of two level surfaces: half-plane and sphere. The red coordinate line is the intersection of two level surfaces: half-plane and cone. The blue coordinate line is the intersection of two level surfaces: sphere and cone. 2. Linear Local Transformations associated with F : scalars and two kinds of vectors We now shift to the Picture A context, where x-space is not necessarily Cartesian. Consider again the possibly non-linear transformation x' = F(x) mapping F: RN→ RN. Imagine a very small neighborhood around the point x in x-space, a "ball" around x. Where the mapping is continuous in both directions, one expects a tiny x-space ball around x to map into a tiny x'-space ball around x' and vice versa. Here is a picture of this situation, where everything in one picture is the mapping of the corresponding thing in the other picture. In particular, we show a small vector in x-space called dx which maps into a small vector in x'-space called dx'. Since F was assumed invertible, it must be invertible locally in these two balls. That is, given a dx above, one can determine dx', and vice versa. Anticipating a few lines below, this means that the matrices S and R will be invertible so neither can have zero determinant. How are these two differential vectors related? For a linear approximation, x'i + dx'i = Fi(x + dx ) ≈ Fi(x) + Σk( ∂Fi(x)/∂xk) dxk => dx'i = Σk( ∂Fi(x)/∂xk) dxk The last line shows an equals sign in the limit that dxk is a vanishing differential. Since Fi(x) = x'i , dx'i = Σk(∂x'i/∂xk) dxk = Σk Rik dxk Rik ≡ (∂x'i/∂xk) Doing the same operation in the other direction gives dxi = Σk( ∂xi/∂x'k) dx'k = Σk Sik dxk' Sik ≡ (∂xi/∂x'k) One can regard Rik and Sik as elements of NxN matrices R and S. In vector notation then, dx' = R(x) dx Rik(x) ≡ (∂x'i/∂xk) R = S-1 // dx'i = Rij dxj dx = S(x') dx' Sik(x') ≡ (∂xi/∂x'k) S = R-1 // dxi = Sij dx'j It is obvious that matrices R and S are inverses of each other, just staring at the above two vector equations. One can verify this fact from the definitions of R and S using the chain rule (RS)ij = Σk RikSkj = Σk (∂x'i/∂xk) (∂xk/∂x'j) = Σk = = δi,j We could get rid of one of these matrices right now, perhaps keeping R and replacing S = R-1, but keeping both simplifies expressions encountered later, so for now both are kept. The letter R does not imply that matrix R is a rotation matrix, although it could be. According to the polar decomposition theorem (Lai p 110), any matrix R (detR ≠ 0) can be uniquely written in the form R = RU = VR where R is a rotation matrix (the same one in RU and VR) and U and V are symmetric positive definite matrices (called right and left stretch tensors) related by U = RTVR. Matrix S could of course be written in a similar manner. Matrices R(x) and S(x') are in general functions of a point in space x' = F(x). As one moves around in space, all the elements of matrices R and S are likely to change. So R and S represent point-dependent linear transformations which are valid for the differentials shown. One might wonder at this point how the vector dx is related to its components dxi and the same question for dx'i and dx'i. As will be shown in Section 6 (f), dx = Σndxn un where the un are x-space axis-aligned basis vectors of the form u1 = (1,0,0,..0) dx' = Σndx'n e'n where the e'n are x'-space axis-aligned basis vectors of the form e'n = (1,0,0,..0) If x-space were Cartesian, one could write un = and e'n = ', but in general the un and e'n vectors do not have (covariant) unit length, as will be demonstrated later. The reader familiar with covariant "up and down" indices will notice that all indices are peacefully sitting "down" in the presentation so far (subscripts, no superscripts). As we carry out our various developmental tasks, that is where all indices shall remain until Section 7, whereupon they will start frantically bobbing up and down, seemingly at will. [ Since rules are made to be violated, we have violated this one in some examples below where non-standard notation would be hard to swallow. ] Are there any "useful objects" that can be constructed from differentials dx and which might then transform according by R or S? The answer is yes, but first we discuss scalars. (a) Scalars A quantity is a scalar with respect to transformation F if it is the same in both spaces. Thus, any constant like π would be a scalar under any transformation. The mass m of a potato would be a constant under transformations that are rotations or translations. A function of space φ(x) is a "field" and it would be a "scalar field" if φ'(x') = φ(x). For example, temperature would be a scalar field under rotations. Notice that φ is evaluated at x, while φ' is evaluated at x' = F(x). As noted in section (k) below, one could be more precise by referring to the objects described here as a "tensorial scalar" and a "tensorial scalar field". (b) Contravariant vectors If transformation F (possibly non-linear) transforms x-space to x'-space without affecting time, then consider the familiar velocity vector, vi = dxi/dt => v = dx/dt Since dt transforms as a constant (scalar) under our selected transformation type, it seems pretty clear that velocity in x'-space can be related to velocity in x-space using the dx' = R(x) dx rule above: v' = R(x) v Even though the matrix R(x) changes as we move around, this linear transformation R is valid at any point x when applied to velocity. Momentum p = mv would work the same way, since mass m is a scalar (Newtonian mechanics). In contrast, unless R(x) is a constant in space (which would be the case only if F were a linear transformation) x' ≠ R(x) x, so in general x itself is not a contravariant vector although dx is. Any vector that transforms according to V' = R(x)V with respect to a transformation F (such as Newtonian velocity and momentum with respect to rotations) is called a contravariant vector. (c) Covariant vectors Much of physics is described by differential equations involving the gradient operator ( the reason for the overbar is given in the next section) i = i = ∂/∂xi which involves an "upside down" differential. Here is how this operator transforms going from x-space to x'-space, again according to the chain rule (implied sum on k) , 'i = 'i = = = Skik = STik k = STik k => ' = ST One can think of as acting on a scalar field φ(x) = φ'(x'), and then the above becomes 'i φ'(x') = φ'(x') = φ(x) = STik k φ(x) => 'φ'(x') = ST φ(x) Since the differential is "upside down", one might expect to transform according to S = R-1 instead of R, but it is really ST that does the job. One could write ' = S in terms of row vectors. Vectors that transform according to V' = ST(x) V such as the gradient operator are called covariant vectors with respect to transformation F. An example of a covariant vector is the electrostatic electric field obtained from the potential Φ = - Φ i = - iΦ = - ∂Φ/∂xi (d) Bar notation In order to distinguish a contravariant from a covariant vector, we shall (for a while) adopt this bar convention: contravariant vectors shall be written V with components Vi and covariant vectors shall be written with components i. This is why overbars were placed on and i and in the previous section. We call this our "developmental notation", as distinct from the Standard Notation introduced in Section 7. The transformation rules for the two vector types can now be written this way: V' = R V contravariant Rik(x) ≡ (∂x'i/∂xk) R = S-1 ' = ST covariant Sik(x') ≡ (∂xi/∂x'k) = STki(x') One could imagine replacing S with some QT to make the second equation more like the first, but of course then RQT = 1 instead of RS = 1. In the Standard Notation, where there are four versions of the matrix R, we shall see that R → Rij and S → Sij = Rji and S can be removed from the picture (see Section 7 (q) ) . (e) Origin of the names contravariant and covariant A justification of the terms covariant and contravariant is presented at the end of Section 7 (t), since the idea is more easily presented there than here. It seems that these terms were first used in 1851 (a half century before special relativity existed) in a paper (see Refs.) by J.J. Sylvester of Sylvester's Law of Inertia fame. Sylvester uses the words covariant and contravariant to describe the relations between a pair of "transformations". In much simpler notation than he uses, if those "transformations" (functions) are F(x) and G(x) and if A is an 3x3 matrix, then the pair F(Ax) and G(Ax) are said to be covariant (or concurrent) the pair F(Ax) and G(A-1x) are said to be contravariant (or reciprocal) The idea is that in comparing the way two things transform, if they both move the same way, then it is covariant, and if they move in opposite directions it is contravariant. In Section 7 (t) this idea is applied to the transformation of two "things", where one thing is the component of a vector like Vn and the other thing is a basis vector onto which a vector is expanded. The connection is a bit distant, but the underlying concept carries through. Notations like y = F(Ax) would have mystified Sylvester in 1851, although in this same paper he introduced two-dimensional arrays of letters and referred to them as "matrices". According to a web piece by John Aldrich of the University of Southampton, J.W. Gibbs in 1881 was the first person to use a single letter to represent a vector (he used Greek letters). It was not until 1901 when his student E.B.Wilson published Gibb's lectures in a Vector Analysis book that the idea was propagated to a wider circle. Wilson converted those Greek letters to bolded ones, The Wilson/Gibbs book was reprinted seven times, the last being 1943. In 1960 it continued as a Dover book and is now available online as a public domain document. (f) Other vector types? Are there any other kinds of vectors with respect to a transformation F? There might be, but only the two types mentioned above are of interest to us in this document. They are both called rank-1 tensors, and there are no other rank-1 tensor types in "tensor analysis" (for rank-n tensors, see Section 7 (j)). Some authors refer to the rank of a tensor as the order of a tensor.) In the Standard Notation introduced later, where contravariant vector components are written with indices up and covariant vectors with indices down, and where the notation is so slick and smooth and automatic, one sometimes imagines there are two kinds of vectors because there are two places to put indices, up and down. It is of course the other way around: the up/down notation was adopted because there are two rank-1 tensor types. Two particular (linear) transformation types of interest are rotations and Lorentz transformations, each of which has a certain number of continuous parameters (3 and 6). As the parameters are allowed to vary over their ranges, the set of transformations can be viewed as elements of a continuous group ( SO(3) and SO(3,1) ). Each of these groups has exactly one "vector representation" ( "1" and "(1/2)(1/2)" ). One should not imagine that somehow the "two-ness" of vector types under general transformations F is connected to there being two vector representations of some particular group. It happens that the Lorentz group does have two "spinor representations" (1/2)0 and 0(1/2), but this has nothing at all to do with our general notion of two kinds of vectors. This subject is discussed in more detail in Section 5 (m). (g) Linear transformations For a linear transformation F, the matrix elements of R and S are constants and don't depend on x or x'. The reason is fairly obvious. For linear x' = F(x) (an added constant would make F non-linear ) F(αx' + βy') = αF(x') + βF(y') => x'i = Fi1x1 + Fi2x2 + .... FiN xN // = Fi(x) where the Fij are constants independent of the coordinates, in which case dx'i = Fi1 dx1 + Fi2 dx2 + .... FiN dxN = Σk Fik dxk so R = F and S = F-1 This is the situation with rotations and Lorentz transformations. (h) Vectors that are contravariant by definition A contravariant vector has been defined above as any N-tuple which transforms the same way that dx transforms with respect to F, namely, dx' = R(x) dx. One might state this as { dx', dx } dx' = R(x) dx contravariant vector Suppose we start with an arbitrary N-tuple V and simply define V' ≡ RV. One would have to conclude that the pair { V', V } transforms as a contravariant vector. { V', V } V' ≡ R(x)V contravariant vector Conversely, one could start with some given V' and define V ≡ S(x)V' (recall S = R-1), and again one would conclude that { V', V } represents a vector that transforms as a contravariant vector. We refer to either process as producing a vector which is "contravariant by definition". Creating a contravariant vector in this fashion is a fine thing to do, as long as the defined vector does not conflict with something that already exists. Example 1: We know that if F is non-linear, the vector x does not transform as a contravariant vector, because x' = R(x)x is not true, where x' = F(x). If we start with x and try to force {x', x} to be "contravariant by definition" by defining x' ≡ R(x) x, this x' conflicts with the existing x' = F(x), so the method of contravariant by definition is unacceptable. Example 2: As another example, consider an N-tuple in x'-space of three masses V' = (m1,m2,m3). The transformation is taken in this example to be regular rotations. Since masses are rotational scalars with respect to such rotations, we know that in an x-space rotated frame of reference we would find V = (m1,m2,m3). We could attempt to set up { V', V } as a vector that is "contravariant by definition" by defining V ≡ SV', but this conflicts with the existing fact that V = (m1,m2,m3), so the method of contravariant by definition is again unacceptable. Example 3: This time F is a general transformation and we start with V' = e'n which are a set of axis-aligned basis vectors in x'-space. We define vectors V = en according to en ≡ Se'n. Then { e'n, en } form a vector which is "contravariant by definition" and e'n = R en (R = S-1). Since the newly defined vector en does not conflict with some already-existing vector in x-space, the method of contravariant by definition in this example is acceptable. This is exactly what is done in the next section with the tangent base vectors en. (i) Vector Fields We considered above vectors like position x (and dx) and velocity v and the vector operator , and we referred to a generic vector as V. Many vectors of interest (in fact, most) are functions of x, which is to say, they are vector fields. Examples are the electric and magnetic fields E(x) and B(x), or the average velocity of a small region of fluid V(x) or a current density J(x). Another example is the transformation F(x). We already mentioned scalar fields, such as temperature T(x) or electrostatic potential Φ(x). The way a scalar temperature field transforms going from x-space to x'-space is this T'(x') = T(x) where x' = F(x) If the transformation is a 3D rotation from frame S to frame S', then T' is the temperature measured in frame S' at point x' and T is the temperature measured at the corresponding point x in frame S and of course there is only one temperature at that point so the numbers are equal. In x'-space one needs the prime on T' because the functional form (how T' depends on the x'i) is not the same as that of T (how T depends on the xi). For example, if transformation F is from 2D Cartesian to polar coordinates, then T'(r,θ) = T(x,y) = T(rcosθ,rsinθ) ≠ T(r,θ) Contravariant and covariant vector fields transform as described above, but now one must show the argument for each field in its own space, and again x' = F(x) : V'(x') = R V(x) contravariant Rik(x) ≡ (∂x'i/∂xk) R = S-1 '(x') = ST (x) covariant Sik(x') ≡ (∂xi/∂x'k) = STki(x') Similar transformation rules apply to tensors of any rank. For example, the metric tensor gab (developmental notation) is a rank-2 contravariant tensor field and the transformation rule is this g'ab(x') = Raa'Rbb'ga'b'(x) or g'ab = Raa'Rbb'ga'b' Often the coordinate dependence of g is suppressed, just as it is for R and S, as shown on the right above. Jumping momentarily into Standard Notation, in special relativity one has x'μ = Λμνxν where F = R = Λ is a linear transformation, and one would then specify the transformation of a contravariant vector field as V'μ(x'α) = Λμν Vν(xα) x'μ = Λμνxν (j) Names and symbols The matrix Rik(x) = (∂x'i/∂xk) is called the Jacobian matrix for the transformation x' = F(x) , while the matrix Sik(x') = (∂xi/∂x'k) is then the Jacobian matrix of the inverse transformation x = F-1(x'). The magnitude of the determinant of Jacobian matrix S will be shown in Section 8 (e) to have a certain significance, and that determinant is called "the Jacobian" = det(S(x')) ≡ J(x'). The author has anguished over what names to give the matrices R and S = R-1. One option was to use R = L, where L stands for the fact that this matrix is describing a Local coordinate system at point x, or a Linearized transformation. But L is always used for differential operators, so that got rejected. R is often called Λ in special relativity, but why go Greek so early? Another option is to use R = J for Jacobian, but J looks too much like "an integer" or angular momentum or "the Jacobian". T for Transformation might have been confused with the tranformation F. Our chosen notation R makes one think perhaps R is a Rotation, but that won't in general be the case. For the moment we will continue to use R and S, where recall RS = 1. The fact that vectors are processed by NxN matrices R and S puts that part of the subject into the field of linear algebra, and that may be the origin of the name tensor algebra as a generalization of this idea (tensors as objects of direct product algebras). Of course the differential calculus aspect of the subject is already highly visible, there are ∂ symbols everywhere (hence the name tensor calculus). (k) Definition of the words "scalar" and "vector". These words have multiple potential definitions. Although we shall lapse very frequently, the following set of definitions would allow for precision statements: A "scalar" is a single number (or expression), a 1-tuple. A "tensorial scalar" is a scalar that transforms under transformation F as a tensorial scalar, which is also known as a rank-0 tensor. An example would be m' = m, mass with respect to 3D rotations. A "vector" is an N-tuple of numbers. A "tensorial vector" is a vector that transforms under transformation F as either a contravariant vector or a covariant vector, so a tensorial vector is a rank-1 tensor. A "scalar field" is a single function of x ( the x-space coordinates). A "tensorial scalar field" is a scalar field that transforms under transformation F as a tensorial scalar field, which is also known as a rank-0 tensor field. For example, f'(x') = f(x) is a scalar field. A "vector field" is an N-tuple of functions of x -- an N-tuple of scalar fields. A "tensorial vector field" is a vector field that transforms under transformation F as either a contravariant vector field or a covariant vector field, so a tensorial vector field is a rank-1 tensor field. The notion of tensor densities described in Appendix D further complicates the nomenclature. One can have scalar densities and vector densities of various weights. 3. Tangent Base Vectors en and Inverse Tangent Base Vectors u'n This entire section is in the context of Picture A, In the previous picture showing dx and dx', one has much freedom to "try out" different differential vectors. For any dx one picks at point x, one gets some dx' according to dx' = R(x) dx. Consider this slightly enhanced version of the previous drawing (red curves added) The point x in x-space (right side) can be regarded as lying on some arbitrary 1-dimensional curve in RN shown on the right in red. Select dx to be the tangent to this curve at point x. That curve will then map into some (probably very different) curve in x'-space which passes through the point x'. The tangent to this curve at the point x' must be dx' = R(x) dx. A similar statement can be made starting instead with an arbitrary curve in x'-space. The tangent dx' there then maps into dx = S(x') dx' in x-space. The curves are in N-dimensional space and are in general non-planar and the tangents are of course N dimensional tangents, so this 2D picture is mildly misleading. We now specialize such that the red curve on the left is a straight line parallel to an x'-space axis, which means the curve on the right is a coordinate line, Admittedly the drawing does not strongly suggest that the red line segment on the left is parallel to an axis in x'-space, but since those axes are not drawn, one cannot complain too strenuously. (a) Definition of the en ; the en are the columns of S First, define a set of N basis vectors in x'-space which point along the positive axes of x'-space, e'n , n = 1,2...N // (e'n)i = δn,i e'1 = (1,0,0...) etc Assume that the dx' arrow above points in this e'n direction so that dx' = e'n dx'n // no implied sum on n where dx'n is a positive differential variation of coordinate x'n along the e'n axis in x'-space. The corresponding dx in x-space will be, dx = S dx' = S [e'n dx'n] = [ Se'n] dx'n ≡ en dx'n where this last equality serves as the definition of en , en ≡ Se'n Vector en = en(x) points along dx in x-space and is tangent to the x'n- coordinate line there at point x. This vector en is generally not a unit vector, hence no hat ^ . Writing the above in components, dx = en dx'n => dxi = (en)i dx'n . But of course dxi = Sin dx'n , and therefore (en)i = Sin = ∂xi/∂x'n or en = ∂x/∂x'n = ∂'nx => (en)i = ∂xi/∂x'n = Sin This says that the vectors en are the columns of the matrix S: S = [e1, e2, e3 .... eN ] matrix = N columns We shall call these en vectors the tangent base vectors. The vectors exist in x-space and point along the various coordinate lines that pass through a point x. If the points on the x'n-coordinate line were labeled with the values of x'n from which they came, one would find that en points in the direction in which those labels increase. As one moves from x to some nearby point, the tangent base vectors all change slightly because in general S = S(x'(x)) and the en = en(x) are the columns of S. Any set of basis vectors which depends on x in this way is called a local basis. In contrast, the corresponding x'-basis e'n shown above with (e'n)i = δn,i is a global basis in x'-space since it is the same at any point x' in x'-space. Since det(S) ≠ 0 due to our assumption that F was invertible, the tangent base vectors are linearly independent and provide a basis for EN. One can of course normalize each of the en to be a unit vector n according to n = en/ |en|. Here is a traditional N=3 picture showing the tangent base vectors pointing along three generic coordinate lines in x-space all of which pass through the point x: Comment on notation. Some authors refer to our en as gn or Rn or other. Later it will be shown that enem = 'nm where 'nm is the covariant metric tensor for x'-space, so admittely this provides a reasonable argument for using gn so that gngm = 'nm. But then the primes don't match which is confusing: the gn are vectors in x-space, while ' is a metric tensor in x'-space. We shall be using yet another g in the form g = det(nm) and a corresponding g'. Due to this proliferation of g objects, we stick with en, the notation used by Margenau and Murphy (p 193). A g-oriented reader can replace e → g as needed anywhere in this document. As for unit vector versions of the en, we use the notation n ≡ en/|en|. Morse and Feshbach use an for this purpose (Vol I p 22). A g-person might use n . A related issue is what symbols to use for the "usual" basis vectors in Cartesian x-space. As noted above, we are using un with (un)i = δn,i as "axis-aligned basis vectors" in x-space. If = 1 for x-space, then these are the usual Cartesian unit vectors (see section (c) below). Many authors use the notation en for these vectors which then conflicts with our use of en as the tangent base vectors. Morse and Feshbach use the symbols i, j, k for our Cartesian u1, u2, u3. Other authors use , , so then un = . Often the notation en is used to represent some generic arbitrary set of basis vectors. For this purpose, we use the notation bn. (b) en as a contravariant vector The situation described above was this, dx' = e'n dx'n x'-space // no implied sum on n dx = en dx'n x-space // no implied sum on n and the full transformation F maps dx into dx'. Since dx is a contravariant vector, the linear transformation R also maps dx into dx'. Thus dx' = R(x) dx e'n dx'n = R(x) en dx'n e'n = R(x) en We can regard the last line as a statement that the vector en transforms as a contravariant vector under F. Written out in components one gets (e'n)i = Rij (en)j δn,i= RijSjn recovering the fact that RS = 1. This is an example of a vector being "contravariant by definition", as discussed in Section 2 (h). These two expansions are easy to show just by verifying that components of both sides are the same: en ≡ Se'n = Σi Sin e'i since (en)j = Σi Sin (e'i)j = Σi Sin δi,j = Sjn = (en)j e'n ≡ Ren = Σi Rin ei since (e'n)j = Σi Rin (ei)j = Σi Rin Sji = (SR)jn = δj,n = (e'n)j (c) a semantic question: unit vectors Above it was noted that e'1 = (1,0,0....). Should this be called "a unit vector" ? It will be seen below that in fact |e'1| = ≠ 1 where ' is the covariant metric tensor in x'-space, and |e'1| is the covariant length of e'1. So e'n is a unit vector in the sense that it has a single 1 in its column vector definition, but it is not a unit vector in the sense that it does not (in general) have unit magnitude (it would if x'-space were Cartesian with g'=1).We take the magnitude = 1 requirement as the proper definition of a unit vector. For this reason, we refer to the e'n in x'-space as just "axis-aligned basis vectors" and they have no "hats". One wonders how such a vector should be depicted in a drawing, see Example 1 (b) below and also Appendix C (e). Example 1: Polar coordinates, tangent base vectors (a) The first step is to compute the matrix Sik(x') ≡ (∂xi/∂x'k) from the inverse equations: x = (x1, x2 ) = (x,y) x' = (x1', x2') = (θ,r) x = F-1(x') ↔ x = rcos(θ) x1 = x2' cos(x1') y = rsin(θ) x2 = x2' sin(x1') So S11 = (∂x/∂θ) = -rsinθ S12 = (∂x/∂r) = cosθ Sik ≡ ( ∂xi/∂x'k) S21 = (∂y/∂θ) = rcosθ S22 = (∂y/∂r) = sinθ S = // det(S) = -r R = S-1 = The tangent base vectors en can be read off as the columns of S e1 = r(-sinθ,cosθ) = eθ = r θ // = r e2 = (cosθ,sinθ) = er = r // = Notice that eθ in this case is not a unit vector. Below is a properly scaled drawing showing the location of the two x'-space basis vectors on the left, and the two tangent base vectors on the right. As just shown, the length of er is 1, while the length of eθ is 2. The tangent base vectors are fairly familiar animals, since er = and eθ = r in usual parlance. If one moves radially outward from point x, the er base vector stays the same, but eθ grows longer. If one moves azimuthally from x to some larger angle θ+Δθ, both vectors stay the same length but they rotate together staying perpendicular. (b) This is a good place to point out that vectors drawn in a non-Cartesian space can have magnitudes which do not equal the length of the drawn arrows. The "graphical arrow length" of a vector v is (vx2 + vy2)1/2, but that is not the right expression for |v| in a non-Cartesian space. For example, as will be shown below, |eθ'| = |eθ| , so the magnitude of the vector e'θ shown on the left above is in fact |eθ'| = r = 2 and not 1, but the graphical length of the arrow is 1 since e'θ = (1,0). See Appendix C (e) for further discussion of this topic with a specific 2D non-orthogonal coordinate system. (c) In this example, two basis vectors e'n in x'-space on the left map into the two en vectors on the right according to en ≡ Se'n. If one were to apply the full mapping x = F-1(x') to each point along the arrows e'n, for some general non-linear F one would find that these arrows map into warped arrows on the right whose bases are tangent to those of the en. Those warped arrows lie on the coordinate lines. For this particular mapping, e'θ maps under F-1 into the warped gray arrow, while e'r maps into er. Example 2: Spherical Coordinates, tangent base vectors x = (x1, x2, x3 ) = (x,y,z) x' = (x1', x2',x3') = (r,θ,φ) x = F-1(x') ↔ x = rsinθcosφ y = rsinθsinφ z = rcosθ S11= (∂x/∂r) = sinθcosφ Sik ≡ (∂xi/∂x'k) S12 = (∂x/∂θ) = rcosθcosφ S13 = (∂x/∂φ) = -rsinθsinφ S21= (∂y/∂r) = sinθsinφ S22 = (∂y/∂θ) = rcosθsinφ S23 = (∂y/∂φ) = rsinθcosφ S31= (∂z/∂r) = cosθ S32 = (∂z/∂θ) = -rsinθ S33 = (∂z/∂φ) = 0 S = R = where Maple computes R as S-1 and finds as well that det(S) = r2 sinθ The tangent base vectors are the columns of S, so er = (sinθcosφ, sinθsinφ,cosθ) |er| = 1 = h'r eθ = r(cosθcosφ,cosθsinφ,-sinθ) |eθ| = r = h'θ eφ = rsinθ(-sinφ,cosφ,0) |eφ| = rsinθ = h'φ and unit vector versions are then r = (sinθcosφ, sinθsinφ,cosθ) = er = θ = (cosθcosφ,cosθsinφ,-sinθ) = eθ = r φ = (-sinφ,cosφ,0) = eφ = rsinθ The unit vectors can be displayed in this standard picture, Notice that (, , ) = (1, 2, 3) form a right-handed coordinate system at the point x = r. (d) The inverse tangent base vectors u'n and inverse coordinate lines A complete swap x' ↔ x for a mapping x' = F(x) of course produces the "inverse mapping". This has the effect of causing R ↔ S in the above discussion. The tangent base vectors for the inverse mapping would then be the columns of matrix R instead of S. We shall denote these inverse tangent base vectors which exist in x'-space by the symbol u'n. Then: (en)i = Sin = ∂xi/∂x'n // the tangent base vectors as above S = [e1, e2, e3 .... eN ] // are the columns of S (u'n)i = Rin = ∂x'i/∂xn // inverse tangent base vectors R = [u'1, u'2, u'3 .... u'N ] // are the columns of R By varying only xn in x-space holding all the other xi = constant, one generates the xn-coordinate lines in x'-space, just the reverse of the earlier discussion of this subject. Then inverse tangent base vectors u'n will then be tangent to these inverse coordinate lines. An example is given below and another in Appendix C. In section (b) above the vector en transformed as a contravariant vector into an axis-aligned basis vector e'n in x'-space e'n = R en (e'n)i = Rij (en )j (en)i = Sin (e'n)i = δn,i The same thing happens here, only in reverse : u'n = S un (u'n)i = Sij (un)j (u'n)i = Rin (un)i = δn,i where now the un are axis-aligned basis vectors in x-space. A prime on an object indicates which space it inhabits. The inverse tangent base vectors u'n are not the same as the reciprocal base vectors En introduced in Section 6 below. Example 1: Polar coordinates: inverse tangent base vectors and inverse coordinate lines It was shown earlier for polar coordinates that, R = S-1 = so the inverse tangent base vectors are given by the columns of R, u'x = ( -sinθ/r,cosθ) // note near θ = 0 that u'x indicates a large negative slope u'y = (cosθ/r,sinθ) // note near θ = 0 that u'y indicates a small positive slope One expects u'x to be tangent to an inverse coordinate line in x'-space which maps to a line in x-space along which only x is varying, which is a horizontal line at fixed y (red). Looking at the small θ region of the plot on the left below, one sees slopes as just described above. For the polar coordinates mapping discussed above, horizontal (red) and vertical (blue) lines in x'-space mapped into circles (red) and rays (blue) in x-space, and the tangent base vectors in x-space were tangent to the coordinate lines there. If one instead takes horizontal (red) and vertical (blue) lines in x-space and maps them back into coordinate lines in x'-space, the picture is a bit more complicated. Since y = rsinθ, the plot of an x-coordinate line (x is varying, y fixed at yi) in x'-space has the form r = yi/sinθ, where yi denotes some selected y value (a red horizontal line), so plotting r = yi/sinθ in x'-space for various values of yi displays a set of inverse x-coordinate lines (red). Similarly r = xi/cosθ gives some y-coordinate lines (blue). Here is a Maple plot: x'-space (θ,r) x-space (x,y) Another example is given in Appendix C. 4. Notions of length, distance and scalar product in Cartesian Space This section can be interpreted in either Picture B or Picture D where the x-space is Cartesian, G=1. Up to this point, we have dealt only with the vector space RN (a vector space is sometimes called a linear space), and have not "endowed" it with a norm, metric or a scalar product. Quantities like dxi above were just little vectors and x + dx was vector addition. Now, for the first time (officially), we discuss length and distance, such as they are in a Cartesian Space, as defined in Section 1. For RN one first defines a norm which determines the "length" of a vector, the first notion of distance in a limited sense. The "usual" norm is the L2 norm given by norm of x = || x || ≡ ( x12 + x22 + .... + xN2 )1/2 ≡ | x | Now we have a normed linear space. One next defines the notion of the distance between two vectors. Although this can be done in many ways, just as there are many possible norms, for RN the "natural metric" is defined in terms of the above L2 norm, so that distance between x and y = metric = d(x,y) ≡ || x - y || = ( [x1-y1]2 + [x2-y2]2 + .... + [xN-yN]2 )1/2 . Now our space is both a normed linear space and a metric space, a combo known as a Banach Space. One finally adds the notion of a scalar product (inner product) in this way (x,y) ≡ Σixiyi ≡ x y // = Σi,j δi,j xi yj which of course implies this special case, (x,x) = x x = Σixi2 = ||x||2 = | x |2 Our space has now ascended to the higher level of being a real Hilbert Space of N dimensions. All this structure is implied by the notation RN, our "Cartesian Space". The length of the vector dx in RN is given by length of dx = distance between vectors x+dx and x ≡ ds ≡ || dx || = To avoid dealing with the square root, one usually writes (ds)2 ≡ || dx ||2 = Σi(dxi)2 = (dx1)2 + (dx2)2 + ... + (dxN)2 = Σi dxi dxi = Σi,j δi,j dxi dxj As shown in the next section, one can interpret δi,j as the metric tensor in Cartesian Space. The cursory discussion of this section is fleshed out in Chapter 2 of Stakgold where the concepts of linear spaces, norms, metrics and inner products are defined with precision. Stakgold compares our N dimensional Cartesian Hilbert Space to the N=∞ dimensional Hilbert Spaces used in functional analysis, where basis vectors might be Legendre polynomials Pn(z) on (-1,1), n = 0,1,2...∞. He has little to say, however, about curvilinear coordinate spaces in this particular book. 5. The Metric Tensor The metric tensor is the heart of the machine of tensor analysis and we shall have a lot to say about it in this section. Each subsection is best presented in the context of one of our Pictures, and there will be some jumping around between pictures. We apologize for this inconvenience and ask forbearance. Hopefully the subsections below will give the reader some experience with typical nitty-gritty manipulations. One advantage of the developmental notation over the standard notation is that matrix methods are easy to use, and they will be used below. We now go to the Picture D context. Comparison with Picture B shows that primes must be placed on objects F, R and S related to the transformation from x-space to x'-space: The various partial derivatives are determined from their definitions, R'ik ≡ (∂x'i/∂xk) R"ik ≡ (∂x"i/∂xk) Rik ≡ (∂x"i/∂x'k) S'ik ≡ (∂xi/∂x'k) S"ik ≡ (∂xi/∂x"k) Sik ≡ (∂x'i/∂x"k) The unprimed S,R can be expressed in terms of the primed objects this way (chain rule) Rik ≡ (∂x"i/∂x'k) = (∂x"i/∂xa) (∂xa/∂x'k) = R"ia S'ak => R = R" S' Sik ≡ (∂x'i/∂x"k) = (∂x'i/∂xa) (∂xa/∂x"k) = R'ia S"ak => S = R' S" (a) Definition of the metric tensor The metric or distance between vectors x and x+dx can be specified as done in Section 4 in terms of the norm of differential vector dx, metric(x+dx, x) = norm( [x+dx] - x) = norm(dx) ≡ ds with the caveat that this is not an official norm, see section (i) below. The squared distance (ds)2 must be a linear combination of products dxidxj just on dimensional grounds. The coefficients in this linear combination form a matrix called the metric tensor (later we show this matrix really is a tensor) (ds)2 = Σi=1N Σj=1N [ metric tensor ]ij dxi dxj This is a bit of chicken and egg because one is really defining "distance" and "metric tensor" at the same time. Each selection of a metric tensor defines the meaning of distance ds in the space of interest. Suppose the length of a small vector dx in a Quasi-Cartesian x-space is known to be ds. Recall from Section 1 that such a space has a diagonal metric tensor G whose diagonal elements are independently either +1 or -1. How might one express this same ds in terms of the other spaces' coordinates x' and x" ? (see Picture D) Going to x'-space one finds, since dx = S'(x') dx', (ds)2 = ΣiGiidxidxi = Σi Gii (ΣkS'ik dx'k) (ΣmS'im dx'm) = ΣkΣm { Σi Gii S'ikS'im } dx'k dx'm Defining the metric tensor in x'-space to be (comment on the bar below) 'km ≡ ΣiGiiS'ikS'im = Σij S'TkiGijS'jm => ' = S'TG S' one then has, with implied summation on the right, (ds)2 = ΣkΣm 'km dx'k dx'm = 'km dx'k dx'm For the transformation from x-space to x"-space in Picture D, a similar result is obtained, "km ≡ Σi GiiS"ikS"im => " = S"TG S" (ds)2 = "km dx"k dx"m Since (ds)2 is a number which is the same in all three systems (that number is the distance between two points in x-space), the quantity 'km dxk' dxm' is a tensorial scalar. The metric tensor is specific to a space; it is a property of the space; it is part of the space's definition. We have placed bars over the g's anticipating what will soon be shown, that these matrices are "covariant" matrices. Then we won't have to go back and fix things up. To summarize, there are three metric tensors for the three spaces in Picture D : = G ' = S'T G S' " = S"T G S" Concerning the invariance of (ds). In the above dicussion, it was assumed that distance (ds)2 is the same in x'-space as it is in x-space. As will be seen soon, this is equivalent to saying that the covariant dot product of any two vectors gives the same number regardless of which space is used to compute the dot product: A B = A' B'. This in turn implies that |A| = |A'| . In other words, it was assumed above that the dot product of two tensorial vectors is a scalar with respect to the underlying transformation F. In our major application, where x-space is Cartesian and x'-space is that of some curvilinear coordinates, it is a requirement that | A | = | A'| . The length of a physical vector is the same no matter how one chooses to describe that vector. Imagine that A is a velocity vector v. The speed |v| of an object is the same number whether one represents v in Cartesian or spherical coordinates. In special relativity one again wants dot products to be scalars and the notion that (ds)2 is a scalar under Lorentz transformations (that is, dxdx = dx'dx') is a key assumption/requirement of the theory. When it is required that (ds)2 ( or A B or |A|) be a tensorial scalar under transformation F from x-space to x'-space, then the metric tensors of the two spaces must be related by ' = S'TG S'. More generally as shown below, if (ds)2 is required the be a tensorial scalar, then one must have ' = ST S where x'-space and x-space are arbitrary spaces with metric tensors ' and . There are, however, applications of transformations where the scalarity of (ds)2 is not required and in fact it is crucial that (ds)2 can change under a transformation. For example, in continuum mechanics one can consider x-space to be a space describing a flow of continuous matter at some initial time t0 and x'-space to be the same flow at a later time t. A general flow has x' = F(x) where x is the position of a continuum "particle" at time t0 and x' is the postion of that same particle at time t. In general F is non-linear. The distance between two differentially spaced particles at the two times is dx and dx', and one has dx' = R dx. The whole point here is that during the flow, the distance vector between two close particles rotates and stretches in some manner, and in general (due to this stretch), |dx| ≠ |dx'| , so (ds)2 is definitely not invariant under the flow (ie, under the transformation F). In this case, the rule ' = ST S does not apply, and one is free to select a metric tensor in each space independently. Since material flows usually occur in Cartesian space, one usually takes g = 1 and g' = 1. In Lai (p 105), the idea that dx' = Rdx translates to dx = FdX, and F is called the deformation gradient and is written F = (x) which is a dyadic like notation discussed in Appendix E and F below. The continuum mechanics application is discussed furthe in Section (o) below. In general, we shall be assuming that in fact (ds)2 is a scalar in almost everything that follows. (b) Inverse of the metric tensor The inverses of the three metric tensors shall be indicates without an overbar, and we shall eventually show these matrices to be "contravariant" matrices and thus deserve no overbar. We thus now define three new g matrices as these inverses, and compute the inverses: g ≡ -1 = G-1 = G // remember G just has +1 and -1 diagonal elements g' ≡ '-1 = (S'T G S')-1 = S'-1 G (S'T)-1 = R' G R'T g" ≡ "-1 = (S"T G S")-1 = S"-1 G (S"T)-1 = R" G R"T Here are the collected facts from above: g = G g' = R'G R'T g" = R" G R"T S = R' S" = G ' = S'TG S' " = S"T G S" R = R" S' g = 1 'g' = 1 "g = 1 Comment: In the Picture C context but with a Quasi-Cartesian x(0)-space, one could take the second column above and write it this way, g = RGRT = STGS g = 1 (ds)2 = km dxk dxm where now the clutter of primes is gone. If x-space is Cartesian so G = 1, then g = RRT and = STS. But we continue with Picture D. (c) A metric tensor is symmetric Any matrix of the form M = ATDA where D is a diagonal matrix (so D=DT) is symmetric: MT = (ATDA)T = ATDA = M // and similarly with A → AT Since all metric tensors shown above match this form, they are all symmetric: gab = gba for any g (with or without an overbar). (d) det(g) and gnn of a Cartesian-generated metric tensor are non-negative If we arrive at x'-space by a transformation F from a Cartesian x-space (as opposed to a Quasi-Cartesian one), we refer to the metric tensor g' in this x'-space as being "Cartesian generated". In this case G = 1 and the metric tensors above are g = RRT and = STS . Any matrix of either of these forms has positive diagonal elements and positive determinant: (ATA)aa = Σb (AT)abAba = Σb (A)baAba = Σb (Aba)2 ≥ 0 // diagonal elements ≥ 0 det(ATA) = det(AT) det(A) = det(A) det(A) = [ det(A) ]2 ≥ 0 // det ≥ 0 To show these results for the AAT form, just replace A→AT everywhere. Recall that transformation F maps RN → RN so the coefficients of the linearized matrices R and S are real, and elements of the metric tensor must therefore also be real. For a Quasi-Cartesian-generated metric tensor, these proofs are invalid since then g = RGRT and = STGS and G ≠1. (e) Definition of two kinds of rank-2 tensors We now switch to Picture A, Recall the vector transformation rules from Section 2 (d), V' = R V contravariant Rik(x) ≡ (∂x'i/∂xk) R = S-1 ' = ST covariant Sik(x') ≡ (∂xi/∂x'k) = STki(x') which can be written out in components V'a = Raa' Va' contravariant Rik(x) ≡ (∂x'i/∂xk) R = S-1 'a = STaa'a' covariant Sik(x') ≡ (∂xi/∂x'k) = STki(x') A rank-1 tensor is defined to be a vector which transforms in one of the two ways shown above. Similarly, a (non-mixed) rank-2 tensor is defined as a matrix which transforms in one of these two ways: M'ab = Raa' Rbb' Ma'b' // contravariant rank-2 tensor 'ab = STaa' STbb' a'b' // covariant rank-2 tensor and again we put a bar over the covariant objects. Digression: Proof that (A-1)T = (AT)-1 for any invertible matrix A: det(A) = det(AT) cof(AT) = [ cof(A)]T since [cof(AT)]ab = cof ( ATab) = cof(Aba) = [cof(A)]ba = [cof(A)]Tab (A-1)T = { [cof(A)]T / det(A) }T = [cof(AT)]T /det(AT) = (AT)-1 This fact is used many times in the manipulations below. (f) Proof that the metric tensor and its inverse are both rank-2 tensors The above rank-2 tensor transformation rules can be written in the following matrix form (something not possible with higher-rank tensors), M' = R M RT // contravariant rank-2 tensor ' = ST S // covariant rank-2 tensor where recall But we now switch these rules to the Picture D context where F maps x'-space to x"-space, M" = R M' RT // contravariant rank-2 tensor " = ST ' S // covariant rank-2 tensor Consider then this sequence of steps: 1 *G * 1 = 1 * G * 1 (S"R") G (S"R")T = (S'R') G (S'R')T // S"R" = 1 S" (R"G R"T) S"T = S'(R'G R'T) S'T // regroup S" g" S"T = S' g' S'T // since g" = R"G R"T and g' = R'G R'T g" S"T = R" S' g' S'T // left multiply by S"-1 = R" g" = R" S' g' S'T R"T // right multiply by S"T,-1 = R"T g" = (R" S') g' (S'T R"T) // regroup g" = (R" S') g' (R" S')T // (AB)T = BTAT g" = R g' RT // expressions in section (b) above for R and S This last result then shows that g' is a contravariant rank-2 tensor with respect to the transformation F taking x'-space to x"-space. Continuing on, g" = R g' RT g"-1 = (R g' RT)-1 g"-1 = ST g'-1 S // RT,-1= ST etc " = ST ' S // ' = g'-1 and this last result shows that ' is a covariant rank-2 tensor with respect to the transformation F taking x'-space to x"-space. This is why we put a bar over this g from the start. These two metric tensor transformation statements can be converted to the Picture A context, g' = R g RT g'ab = Raa'Rbb'ga'b' // g is a contravariant rank-2 tensor ' = ST S 'ab = STaa'STbb'a'b' // is a covariant rank-2 tensor Since RS = 1, the equations can be inverted to get g = S g' ST gab = Saa'Sbb'g'a'b' = RT ' R ab = RTaa'RTbb''a'b' Two more useful variations of the above are Rg = g' ST Rabgbc = g'abScb S = RT ' abSbc = Rba'bc (g) Metric tensor converts vector types We continue in Picture A. Suppose V is a contravariant vector so V' = RV. Construct a new vector W with the following properties ( see Section 7 (u) concerning "covariant equations") W = V x-space W' = ' V' x'-space Is vector W one of our two vector types, or is it neither? One must examine how it transforms under F: W' = ' V' = (ST S) (RV) = ST (SR)V = ST V = ST W Therefore this new vector W is a covariant vector under F, so it should have an overbar, ≡ V This covariant vector can be regarded as the covariant partner of contravariant vector V. This shows the general idea that applying to any contravariant vector produces a covariant vector! So this is one way to construct covariant vectors if we have a supply of contravariant ones. Conversely, starting with a known covariant vector , one can construct a contravariant vector V ≡ g . Thus, every vector of either type can be thought of as having a partner vector of the other type. An obvious notation is to write as so no extra letter is needed. Then one has = V V = g i = ij Vj Vi= gijj (h) Vectors in Cartesian space Theorem: There is no distinction between a contravariant and a covariant vector in Cartesian space. Proof: Pick a contravariant vector V. Since = 1, ≡ V = V . But is a covariant vector. Since = V , every contravariant vector is also covariant and vice versa. In other words, if g = 1, every vector is the same as its covariant partner vector. The transformation rules in this case are V' = R V ' = ST = ST V Although the vectors V and are the same, eliminating V shows that ' and V' are not the same. One finds that ' = (STS) V' = ' V', so ' = ' V' ≠ V'. (i) Metric tensor: covariant scalar product and norm For a Cartesian space, Section 4 defined the norm as the length of a vector, the metric as the distance between two vectors, and the scalar product (inner product) as the projection of one vector on another. The official definitions of norm, metric and scalar product require non-negativity: | x | ≥ 0, d(x,y) ≥ 0, and x x ≥ 0. For non-Cartesian spaces, the logical extensions of these three concepts can result in all three quantities being negative. Nevertheless, we shall use the term "covariant scalar product" with notation A B as defined below, as well as the notation |A|2 ≡ A A where |A| will be called the length, magnitude or norm of A, even though these objects are not true scalar products or norms. In the curvilinear application of tensor analysis, where x-space is Cartesian, since the norm and scalar product are tensorial scalars, and since they are non-negative in Cartesian x-space, the problem of negative norms does not arise in either space. How do authors handle this problem? Some authors refer to A A as "the norm" of A (e.g., Messiah bottom p 878 discussing special relativity), which is our |A|2. For a general 4-vector A in special or general relativity, most authors just write A A (AμAμ in standard notation), note that the quantity is invariant under transformations, but don't give it a name. Whereas we use the bold for this covariant dot product, most special relativity authors prefer to reserve this bold dot for a 3D spatial dot product, and then the 4D dot product is written with some "less bold dot" such as A.B or A•B. Typical usage then in standard notation would be p•p = pμpμ = p02 - pp (see for example Bjorken and Drell p 281). Without further ado, we define the "covariant scalar product" of two contravariant vectors (a new and different use of the word "covariant", but the same as appears in Section 7 (u) ) as: A B ≡ abAaBb = abBbAa = baBaAb = abBaAb = B A This covariant scalar (or dot) product is more interesting and useful than the object AaBa because the covariant scalar product of two contravariant vectors is a tensorial scalar, as we now show (Picture A) A' B' = 'abA'aB'b = 'ab(Raa'Aa') (Rbb'Bb') = 'ab Raa' Rbb' Aa' Bb' = [ (RT)a'a 'ab Rbb' ] Aa' Bb' = [RT ' R]a'b' Aa' Bb' = a'b' Aa' Bb' = ab Aa Bb = A B Recall that for any contravariant vector B, there is a partner covariant vector a = abBb. Using this partner one can restate the above covariant scalar product as A B = abAaBb = Aa a or, taking instead b = baAa , A B = b Bb = a Ba And finally, if in A B = Aa a we write Aa = gabb , we get A B = Aa a = gabba = gabab where the scalar product is now expressed in terms of the covariant partner vectors and . To summarize, there are four different ways to write this covariant scalar product : A B = abAaBb = Aa a = a Ba = gabab = B A Using the appropriate expressions on the above line, one may conclude that the covariant dot product of any two tensorial vectors is a tensorial scalar. In the special case that A = B, we use the shorthand notation (with caveat as noted above) |A|2 ≡ A A Going back to the a result of section (a), one sees the (ds)2 distance squared in a new light, (ds)2 = 'km dx'k dx'm = dx' dx' = dx dx = a scalar with respect to F so ds is sometimes called "the invariant distance". In special relativity, using the Bjorken and Drell notation noted above where g'μν = diag(1,-1,-,1,-1) and c=1, one writes ( standard notation) (dτ)2 = g'μν dx'μdx'ν = dxμdxμ = dx'• dx' = dx • dx = a Lorentz scalar = (dt)2 - dx dx , xμ = (t,x) and dτ is called "the proper time", a particular case of the invariant distance ds. Notice that (dτ)2 < 0 for a spacelike 4-vector dxμ, meaning one that lies outside the future and past lightcones (|dx| > |dt| ). We now restore to our covariant definition. Going back to Section 3 and the vectors e'n and en, a claim made there can now be verified: |e'n|2 = e'n e'n = en en = |en|2 => |e'n| = |en| (j) Metric tensor and tangent base vectors The context of Picture A continues, Recall this fact from Section 3, S = [e1, e2, e3 .... eN ] where the columns of S are the tangent base vectors. It follows that (see end of section (h)) ' = ST S = [e1, e2, e3 .... eN ]T [e1, e2, e3 .... eN ] so e1e1 e1 e2 e1 e3 ...... e1 eN e2e1 e2 e2 e2 e3 ...... e2 eN ' = e3e1 e3 e2 e3 e3 ...... e3 eN ........ eNe1 eN e2 eN e3 ...... eN eN since, enT em = (en)i ij (em)j = ij(en)i(em)j = en em using the covariant scalar product defined in the previous section. Taking the n,m component of the above matrix equation, one gets 'mn = em en or 'mn = ∂'mx ∂'nx which makes a direct connection between the covariant metric tensor in x'-space and the tangent base vectors en in x-space. A less graphical derivation of this fact is 'nm = (STS)nm = STna ab Sbm = ab San Sbm = ab (en)a(eb)n ≡ en em . Therefore, the tangent base vectors will only be mutually orthogonal when the metric tensor ' of x'-space is a diagonal matrix. We refer to the coordinates of an x'-space having a diagonal metric tensor as comprising an orthogonal coordinate system. At any point x in x-space, the tangents en to the N coordinate lines passing through that point are orthogonal. Most examples below will involve such systems, with Appendix C providing a non-orthogonal example. In particular, 'mn = em en lets us write the length of a tangent base vector in terms of the corresponding diagonal element of ', |en|2 = en en = 'nn => |en| = => n = en / The quantities |en| = are called scale factors and are sometimes written h'n or Q'n or H'n. h'n ≡ Q'n ≡ |en| = As a reminder, had we called x'-space something like ξ-space, there would be no primes on these symbols, but then if ≠1 there would be confusion as to which space the symbols applied. Section (d) above showed that 'nn ≥ 0 when x-space is Cartesian. This is the usual case for the curvilinear coordinates application, and so in this case the scale factors h'n are always real and positive. Note: Some authors refer to the scale factors h'n as the Lamé coefficients, while other authors refer to Rij as the Lamé coefficients which they call hji. (Lame) (k) The Jacobian J The context of Picture A continues, First of all, note that since RS = 1, det(S) = 1/det(R) The Jacobian J(x') is defined as follows, J(x') ≡ det(S(x')) = det(∂xi/∂x'k) = 1/det(R(x(x')) = 1/ det(∂x'i/∂xk) Note 1: Objects which relate to the transformation between x-space and x'-space cannot themselves be tensors because tensor objects must be associated with a specific space, the way V(x) is a vector in x-space and V'(x') is a vector in x'-space. Thus Sij(x') = ∂xi/∂x'k , although a matrix, is not a rank-2 tensor. Similarly, J(x'), while a "scalar" function, is not a rank-0 tensorial scalar. One does not ask how S and J themselves "transform" in going from x-space to x'-space. Note 2: An alternative notation used by some authors is this J(x,x') ≡ det(S(x,x')) = det(∂xi/∂x'k) as if x and x' were independent variables. In our presentation, x' = F(x) is not an independent variable but is determined by F(x). Just as one might write f'(x') = ∂f/∂x', we write J(x') = det(∂xi/∂x'k). The connection would be J(x') = J(x=F-1(x'),x')) = J(x(x'),x'). Note 3: Other sources often use the notation | M | to indicate the determinant of a matrix. We shall use the notation det(M), and reserve | | to indicate the magnitude of some quantity, such as |J| below. The determinant of any NxN matrix S may be written (εabc.. is the permutation tensor, Section 7 (h)), det(S) = εabc...x Sa1 Sb2 ... SxN For our particular S with Sin = (en)i this becomes det(S) = εabc...x (e1)a(e2)b....... (eN)x so J is related to the tangent base vectors by J = εabc...x (e1)a(e2)b....... (eN)x . It was shown in section (f) that ' = ST S and g' = R g RT , these being the transformation rules for covariant and contravariant rank-2-tensors. Therefore det(') = det(STS) = det(ST)det()det(S) = det(S)det(S)det() = J2 det() det(g') = det(RgRT) = det(R)det(g)det(RT) = det(R)det(R)det(g) = J-2 det(g) or det(') = J2 det() => J2 = det(') / det() = [det(S)]2 det(g') = J-2det(g) It is a tradition to define certain scalar (but not tensorial scalar) objects with the same name g and g', g(x) ≡ det((x)) = 1/det(g(x)) // in x-space g'(x') ≡ det('(x')) = 1/det(g'(x')) // in x'-space So that J2(x') = det('(x')) / det((x)) = g'(x') / g(x) Normally the argument dependence is suppressed and one then writes J2 = det(')/ det() = g'/g As explained in Appendix D (a), the equation g' = J2 g says that g, instead of being a tensorial scalar, is a scalar density of weight -2. One must be a little careful to distinguish the scalars g and g' from the tensors gij and g'ij expressed in matrix notation as g and g'. In taking the square root of the above equation, one gets (g'/g)1/2 = ± J so a choice is required. Sometimes a useful choice is to take the + sign regardless of the sign of J, which means that when J < 0, one is taking the negative branch of the square root. It is convenient to make the following definition, called the signature of the metric tensor, s = sign[det()] Since g = 1, one has det(g)det() = 1 so that sign[det()] = sign[det(g)] . Since det(') / det() = [det(S)]2, one has sign[det(')] = [det()]. Therefore: s = sign[det()] = sign[det(g)] = sign[det(')] = sign[det(g')] = sign(g) = sign(g') Since transformation F is assumed invertible in its domain and range, one cannot have det(S) = 0 anywhere except perhaps on a boundary. Since det(') = [det(S)]2det(), if we assume det() vanishes nowhere in the x-space domain of F, then det(') ≠0 everywhere in the range of F. The conclusion with this assumption is that the signature s is always well-defined. Obviously, the quantities sg and sg' are both positive, and since J2 = g'/g one can write |J| = / = | det(S) | = For the curvilinear coordinates application, x-space is Cartesian, det() = 1, and thus s = 1 and then |J| = = | det(S) | // curvilinear For the relativity application, x-space is Minkowski space with det() = -1 so s = -1 and |J| = = | det(S) | // relativity Here then is a summary of the results of this section: J(x') ≡ det(S(x')) = det(∂xi/∂x'k) = 1/det(R(x(x')) = 1/ det(∂x'i/∂xk) g ≡ det() g' ≡ det(') g' = J2g => g is a tensor density of weight -2 s ≡ sign[det()] = sign[det(g)] = sign[det(')] = sign[det(g')] = sign(g) = sign(g') |J| = / = | det(S) | = Note: Weinberg p 98 (4.4.1) defines g = -det(gij). This is the only one of Weinberg's conventions that we have not adopted, so in this paper it is always true that g ≡ + det(gij) even though this is -1 in the application to special relativity. Carl Gustav Jacob Jacobi (1804 –1851). German, Berlin PhD 1825 then went to Konigsberg, did much in a short life. Elucidated the whole world of elliptic integrals and functions, such as F(x,k) and sn(x;k), which occur even in simple problems like the 2D pendulum. Wiki claims he promoted Legendre's ∂ symbol for partial derivatives (used throughout this document) and made it a standard. Among many other contributions, he saw the significance of the object J which now bears his name: "the Jacobian". The Jacobi Identity is another familiar item, a rule for non-commuting operators [x,[y,z]] + [z,[x,y]] + [y,[z,x]] = 0 which finds use with quantum mechanical operators and matrices, and more generally with Lie group generators. (l) Some relations between g, R and S in Picture C In Picture C, the statement of the rank-2 tensor transformation of g and becomes g = RRT = STS which can be written in a variety of ways, RT = (SR)RT = S(RRT) = S g => R = g ST => 1 = S g ST ST = ST(RTST) = (STS)R = R => S = RT => 1 = RT R In summary: g = RRT RT = S g R = g ST 1 = S g ST = STS ST = R S = RT 1 = RT R The diagonal elements of and g are given by nn = Σn STniSin = Σn (Sin2) = Σi (∂xi/∂x'n)2 gnn = Σn RniRTin = Σn (Rni2) = Σi (∂x'n/∂xi)2 If the x'i are orthogonal coordinates, then nm = h'n2δnm and gnm = h'n-2δnm where the h'n are the scale factors mentioned above in section (j). These scale factors may then be expressed as ( M&F p 23 1.3.4) h'n2 = nn = Σi (∂xi/∂x'n)2 h'n-2 = gnn = Σi (∂x'n/∂xi)2 Example 1: Polar coordinates: metric tensor and Jacobian Picture C continues (so now θ = x1 and r = x2) and the metric tensor for polar coordinates will be computed in two ways. On the last visit to this example ( end of Section 3), it was shown that S = = [ e1, e2 ] e1 = r(-sinθ, cosθ) e2 = (cosθ, sinθ) One way to compute is this: ( 1=θ, 2=r) = STS = = => θθ = r2 rr = 1 Another way is this: = = // det() = r2 Notice that this metric tensor is in fact symmetric, and that one of its elements is a function of the coordinates. The length2 of a small vector dx can be written (ds)2 = km dxk dxm = θθ dθ dθ + rr dr dr = r2 (dθ)2 + (dr)2 The Jacobian is given by J(r,θ) = det(S(r,θ)) = det = -r so |J| = r and g = J2 = r2, = r Example 2: Spherical coordinates: metric tensor and Jacobian As with Example 1, Picture C is used, wherein (x1, x2, x3) = (r,θ,φ) . In our last visit to this example (end of Section 3) it was found that S = The metric tensor is then given by Maple as = STS = det() = r4sin2θ so that 11= rr = 1 h1 = hr = = 1 22= θθ = r2 h2 = hθ = = r 33= φφ = r2sin2θ h3 = hφ = = rsinθ The Jacobian is found by Maple to be, J(r,θ,φ) = det(S) = r2sinθ Differential distance is then (ds)2 = km dxk dxm = (dr)2 + r2(dθ)2 + r2sin2(dφ)2 and if dφ = 0, this agrees with the polar coordinates result. (m) Special Relativity and its Metric Tensor: vectors and spinors In this section the Standard Notation introduced below in Section 7 is used. In that notation Rij is written Rij , contravariant vectors Vi are written Vi, and covariant vectors j are written Vj. It is a tradition in special and general relativity to use Greek letters for 4-vector indices and Latin letters for spatial 3-vector indices. The (Quasi-Cartesian) metric tensor of special relativity is frequently taken as G = diag(1,-1,-1,-1) and the ordering of 4-vectors as xμ = (t,x,y,z) where c=1 (speed of light) and μ= 0,1,2,3 ( Bjorken and Drell p 281). General relativity people often use G = diag(-1,1,1,1) ≡ η instead (Weinberg p 26). Still other authors use G = 1 and xμ = (it,x,y,z) where i is the imaginary i, but this approach does not easily fit into our tensor framework which is based on real numbers. A Lorentz transformation is a linear transformation Fμν x'μ = Fμν xν = Rμν xν => xν is a contravariant vector and a theory requirement is that invariant length be preserved x'.x' = x.x = scalar. Special relativity also requires that the metric tensor G be the same in all frames, since no frame is special, so G' = G. But this says, in our old notation, that R G RT = G. This condition restricts the (proper) Lorentz transformations to be rotations, boosts (velocity transformations), or any combination of the two. In particular, R G RT = G => det(R G RT) = det(G) => det(R) det(G) det(RT) = det(G) => [det(R)]2 (-1) = (-1) => det(R) = ±1 Proper Lorentz transformations have det(R) = det(F) = +1, and here are two examples. First, a boost transformation in the x direction, Fμv = = exp(-ibK1) where (K1)μν = and second, a rotation transformation about the x axis, Fμv = = exp(-irJ1) where (J1)μν = The matrices K1 and J1 are called generators and are a part of a set of six 4x4 matrices Ji and Ki for i = 1,2,3. These 6 generator matrices satisfy a set of commutation relations known as a Lie Algebra, [ Ji, Jj] = +i εijkJk // [ A,B ] ≡ AB - BA [ Ji, Kj] = +i εijk Kk [ Ki, Kj] = -i εijkJk In these commutators, the generators Ji and Ki can be regarded as abstract non-commuting operators, while the specific 4x4 matrices shown above for J1 and K1 are just a "representation" of these abstract operators as 4x4 matrices. The six 4x4 generator matrices Ji and Ki are (g = G = diag(1,-1,-1,-1) ) (Jμν)αβ = i ( gμαδνβ – gναδμβ) (J1)αβ ≡ (J23)αβ = i ( g2αδ3β – g3αδ2β) and cyclic 123 (K1)αβ ≡ (J01)αβ = i ( g0αδ1β – g1αδ0β) and cyclic 123 where (-i)(Jμν)αβ = ( gμαgνβ – gναgμβ) is a rank-4 tensor, antisymmetric under μ↔ν and α ↔ β. An arbitrary Lorentz transformation can be represented as Fμv(r,b) = [ exp {– i ( rJ + bK)} ]μν where the 6 numbers r and b are called parameters (rotation and boost) and this F is a combined boost/rotation transformation (note that eA+B ≠ eAeB for non-commuting matrices A,B). The product of two such Lorentz transformations is also a Lorentz transformation, and in fact the transformations form a continuous group known as the Lorentz Group, which then has 6 parameters. The first two commutators shown above ( all J and all K) are each associated with a 3 parameter continuous group called the rotation group. The abstract generators of this group can be "represented" as matrices of any dimension, and are labeled by a number j such that 2j+1 is the matrix dimension. For example, the 2x2 matrix representation of the rotation group is labeled by j = 1/2, and is called the spinor representation and is associated in physics with the "intrinsic spin" of particles of spin 1/2 such as electrons. The vectors (spinors) in this case have two elements, and (1,0) and (0,1) are "up" and "down". Representations of the Lorentz group have labels {j1, j2}, where j1 is for the J-generated rotation subgroup, and j2 for the K-generated rotation subgroup, and are usually denoted j1j2. Such a representation then has vectors containing (2j1+1)(2j2+1) elements. In the case 1/21/2 there are 2*2=4 elements in a vector, and when these elements are linearly combined in a certain manner, they form the 4-vector object which one writes as Aμ such as xμ. This is the "vector representation" of the Lorentz group upon which is built the entire edifice of special relativity tensor algebra. One can also consider two other representations of the Lorentz group which are pretty obvious: 1/2 0 and 1/2 0. These are 2x2 matrix representations and they are different 2x2 representations. For each representation one can construct a whole tensor analysis edifice based on 2-vectors. Just as with the 4-vectors, one has contravariant and covariant 2-vectors.The two representations 1/2 0 and 1/2 0 are called spinor representations since they are each 2-dimensional. Since there are two distinct spinor representations, one needs some way of distinguishing them from each other. One representation might be called "undotted" and the other "dotted" and then there are four 2-vector types to worry about, which transform this way V'a = RabVb V' = RV contravariant 2-vectors V'a = RabVb V' = RV covariant 2-vectors where now dots on the indices indicate which Lorentz group representation that index belongs to. The 2x2 matrices Rab and R are not the same. A typical rank-2 tensor would transform this way, X' a = Raa'R' Xa'' This then is the subject of what is sometimes called Spinor Algebra as opposed to Tensor Algebra, but it is really just regular tensor algebra with respect to the two spinor representations of the Lorentz group. We have inserted this blatant digression just to show that the general subject of tensor analysis includes all this spinor stuff under its general umbrella. In closing, Maple shows that the metric tensor G is indeed preserved under boosts and rotations. In Maple evalm(Bx &* G &* transpose(Bx)) means Bx G BxT and Maple is just verifying that Bx G BxT = G and similarly Rx G RxT = G : (n) General Relativity and its Metric Tensor In general relativity a Picture of interest is Picture C but the x(0) space is replaced by a Quasi-Cartesian space with coordinates ξi with the metric tensor of special relativity. This ξ-space represents a "freely-falling" coordinate system in which the laws of special relativity apply and the metric tensor is taken to be G = diag(-1,1,1,1) ≡ η. The xi are the coordinates of some other coordinate system. There is some transformation x = F(ξ) which defines the relationship between these two systems. The covariant metric tensor in x-space is written = STGS = STηS. Using the Standard Notation introduced in Section 7 below, this is usually written as = STηS // result from Section 5 (b) gdn = ST ηdnS // where (ηdn)ab = ηab gμν = (ST)μα ηαβ Sβν = Sαμ ηαβ Sβν // Standard Notation as in Section 7 gμν = (∂ξα/∂xμ) ηαβ (∂ξβ/∂xν) gμν = (∂ξα/∂xμ) (∂ξβ/∂xν) ηαβ // p 71 (3.2.7) The last line then defines the gravitational metric tensor in x-space based on the transformation ξ = F-1(x) = ξ(x). ( This and the following references are from the book of Weinberg, see References.) Newton's Second Law ma = F appears this way in general relativity m (∂2xμ/∂τ2) = Fμ - m Γμνλ (∂xν/∂τ) (∂xλ/∂τ) // p 123 (5.1.11 following) where Fμ is an externally applied force, but there is then an extra bilinear velocity-dependent term which represents an effective gravitational force (it acts on mass m) arising from spacetime itself. The object Γμνλ is called the affine connection and is related to the metric tensor in this way. Γμνλ = ½ gμσ( ∂νgλσ + ∂λgνσ – ∂σgνλ ) // p 75 (3.3.7) These brief comments are only meant to convince the reader that the equations of general relativity also have their place under the general umbrella of tensor analysis as discussed in this document. The fact that Γμνλ is not a mixed rank-3 tensor is discussed in Section 7 (v). (o) Continuum Mechanics and its Metric Tensors One can describe (Lai) the forward "flow" of a continuous blob of matter by x = x(X,t) where X = x(X,t0). A "particle" of matter (imagine a tiny cube) that starts at location X at time t0 ends up at x at time t. Two points in the flow separated by dX at t0 end up separated by some dx at t. The relation between them is given by dx = F dX where F is called the deformation gradient (a rank-2 tensor). F describes how a particle starting say with a cubic shape at t0 gets deformed into some parallelepiped (3-piped) shape at t. The finite-time flow x = x(X,t) from time t0 to time t can be thought of as a generic (generally non-linear) transformation of the form x = F(X) as in Section 1 above (but we replace our usual F by F to avoid confusion between two F symbols: F is now the linearization of transformation F at a point x). The two times are regarded as fixed parameters. Both the starting X-space and the ending x-space are Cartesian spaces, since this flow occurs in the physical world! Thus, the metric tensors for x-space and X-space are both 1 for Cartesian coordinates in each of these spaces. This in turn means that the covariant tensor analysis concepts such as the preservation of the length of a vector under the transformation go out the window, and in fact the vector dX typically gets stretched as dX → dx so that | dX | ≠ | dx |. In order to put this flow into the notation of this document, let X → x and x → x' so that continuum mechanics this document (Forward Flow) x, X ↔ x', x x = x(X,t) ↔ x' = F(x) // Lai p70 (3.1.4) dx = F dX ↔ dx' = R dx // as in Section 2 // Lai p86 (3.7.6), p105 (3.18.3) F ↔ R X = Cartesian ↔ = 1 x = Cartesian ↔ ' = 1 B = FFT ↔ ' = RRT // as in Section 5 (l) // Lai p121 (3.25.2) Thus, the deformation gradient F is just the R matrix of the forward transformation x = x(X,t) = F(X). What we might call the "would-be" metric tensor, ' = RRT = STS ( that is, the ' metric tensor that would have resulted in scalars being true scalars under the transformation), appears as B = FFT and this is known as the left Cauchy-Green deformation tensor (manifestly symmetric). Regarding the above as a description of forward flow, one could consider instead the inverse flow process, but with F having the same meaning as in the forward flow, dx = F dX. Then the inverse flow translation table would be this : continuum mechanics this document (Inverse Flow) X, x ↔ x', x X = X(x,t) ↔ x' = F(x) dX = F-1 dx ↔ dx' = R dx // as in Section 2 x = F(X) F-1 ↔ R F ↔ S // S = R-1 X = Cartesian ↔ = 1 x = Cartesian ↔ ' = 1 C = FTF ↔ ' = STS // as in Section 5 (l) // Lai p114 (3.23.2) In this direction the would-be ' tensor corresponds to C = FTF which is the right Cauchy-Green deformation tensor. Given the above flow situation, it is then possible to add two more transformations F1 and F2 which take X-space and x-space to independent sets of curvilinear coordinates X' and x': and we then have an interesting triple application of the notions of Section 1 to a real-world problem. This drawing is the implicit subject of Section 3.29 (p131) of Lai. In (reverse) dyadic notation the deformation gradient is written F = (x) where means (X)so that dx = F dX = (x) dX Fij = (x)ij = ∂j(X)xi = ∂xi/∂Xj The (x) notation is explained in Appendix E, and in Appendix F the object (v) for an arbitrary vector field v(x) is expressed in general curvilinear coordinates. Section 8 discusses how length, area and volume transform under a transformation like F. In that discussion we can regard the Section 8 picture with its "Cartesian-View" x'-space and the skewed N-piped to its right as describing the (inverse) fluid flow situation for a tiny fluid particle. It is shown there that the length, area and volume magnitude ratios are given by (converted to developmental notation), | dx(n)|/ dL'n = h'n = ['nn]1/2 = the scale factor for edge dx(n) | dA(n)|/ dA'n = (1/h'n) |J| = (1/h'n) g'1/2 = ['nn g']1/2 = [cof('nn)]1/2 |dV| / dV' = |J| = g'1/2 // g' ≡ det('ij) = J2 , ' = STS which can be translated into our inverse flow context as follows : | dx(n)| / | dX(n)| = ['nn]1/2 = [(FTF)nn]1/2 = [Cnn]1/2 // Lai p114 (3.23.6) | dan| / | dAn| = [cof('nn)]1/2 = [cof((FTF)nn)]1/2 = ['nn g']1/2 // Lai p129 (3.27.11)* |dv| / |dV| = |J| = [det('ij)]1/2 = [det(FTF)]1/2 = |det(F)| // Lai p 130 (3.28.3) where edge area volume X-space : dX(n) dAn dV time t0 x-space : dx(n) dan dv time t Thus, for example, the volume change of a "flowing" particle of continuous matter is given by the Jacobian |J| = |detF| associated with the deformation gradient tensor F. We put quotes on "flowing" only because this might be a particle of steel that is momentarily flowing a very small amount during an oscillation or in response to an applied stress. * Lai's area ratio is presented in the following unusual manner: | dan| / | dAn| = detF | F-1,Tun| = detS |RTun| where un is our usual Cartesian space unit vector. But detS = J = g'1/2 (Section 5 (k)) and [RTun]i[RTun]i = Rni Rni = (RRT)nn = 'nn => |RTun| = ['nn]1/2 which gives | dan|/ | dAn| = detF | F-1,Tun| = g'1/2 ['nn]1/2 = ['nn g']1/2 in agreement with the value given above. 6. Reciprocal Base Vectors En and Inverse Reciprocal Base Vectors U'n This entire Section uses the Picture A context, (a) Definition of the En Although various definitions are possible, we shall define the reciprocal tangent vectors En in the following manner (implied sum on i) En ≡ g'ni ei = g'ni∂'ix => en = 'niEi // since ' = g'-1 Comment: Notice how this differs in structure from the rule for forming a covariant vector from a contravariant one, n = 'niVi In the last equation, the right side is a linear combination of vector components Vi, while in the previous equation the right side is a linear combination of vectors ei. In this case, i is a label on ei , whereas in the other case i is an index on Vi. Labels and indices are different. Since the tangent base vectors ei are contravariant vectors in x-space (Section 3 (b)), and since En is a linear combination of the ei, the En are also contravariant vectors in x-space. Notice that in the definition En ≡ g'ni ei, these two x-space vectors are related by the metric tensor of the other space. One can express the components of En in two ways, (En)k ≡ g'ni (ei)k = Skig'ni / since (ei)k ≡ Ski, Section 3 = g'niSki = Rnigik // since Rabgbc = g'abScb, end of Section 5 (f) so that (En)i = g'naSia = giaRna // sum on second indices Applying R to both sides of En ≡ g'ni ei gives En transformed into x'-space, E'n = g'nk e'k so (E'n)i = g'nk (e'k)i = g'nkδk,i = g'ni (b) The Dot Products and Reciprocity (Duality) Three covariant dot products are of great interest. The first is this (Section 5 (f) for last step) en em = ij (en)i (em)j = ij Sin Sjn = STni ij Sjm = ( ST S)nm = 'nm The second is En em = ij (En)i (em)j = ij gia Rna Sjm = δj,a Rna Sjm = Rnj Sjm = (RS)nm = δn,m and the third is En Em = ij (En)i (Em)j = ij gia Rna gjb Rmb = δj,a Rna gjb Rmb = Rnj gjb Rmb = Rnj gjb RTbm = (R g RT)nm = g'nm To summarize, en em = 'nm En em = δn,m En Em = g'nm Using the transformed E'n defined above, one finds that E'n e'm = 'ij (E'n)i(e'm)j = 'ij g'ni δm,j = 'im g'ni = δn,m which is consistent with the fact that this is a covariant dot product of tensorial vectors: E'n e'm = En em = δn,m The other two dot products above work this same way, so e'n e'm = 'nm E'n e'm = δn,m E'n E'm = g'nm Notes on Reciprocity (Duality) 1. The reciprocal base vectors En are more usually defined as being those vectors which satisfy the equations En em = δn,m where the em are known. Each En vector has N components so the full set of En vectors has N2 components. As n and m take all values, En em = δn,m represents N2 linear equations. A solution exists since S is invertible (the em form a complete set). The solution is unique and in fact gives our assumed definition above En ≡ g'ni ei . Here is a fast solution of this Cramer's Rule problem using matrix notation: En em = δn,m => ij(En)i(em)j = δn,m => (En)i ij Sjm = δn,m Let Ani ≡ (En)i . Then have A S = 1, => A = ( S)-1 = S-1 -1 = R g = g' ST (end Sec 5f). Therefore A = g' ST => (En)i = Ani = g'nj (ST)ji = g'nj (ej)i => En = g'nj ej . QED 2. In general, if one has An am = δn,m, the vectors Am are said to be "reciprocal" to the am and vice versa, so the vectors En are reciprocal to the tangent base vectors en. 3. Some authors refer to An am = δn,m as a "duality relation" and either set of vectors is "dual to" the other set. The En are referred to as the dual vectors to en. 4. If the am are contravariant vectors, then An will also be contravariant and then An am is a tensorial scalar. Therefore if An am = δn,m, then so also A'n a'm = δn,m in x'-space, where am' = Ram and A'n = RAn. For example, E'n e'm = δn,m in x'-space where em' = Rem and E'n = REn . 5. In section (e) we shall encounter another dual pair Un um = U'n u'm = δn,m which is associated with the inverse transformation x = F-1(x'). 6. One major significance of the equation An am = δn,m is that it allows the following expansions: V = Σn kn An where km = V am V = Σn cn an where cm = V Am so that for example am V = am [Σn kn An] = Σn kn am An = Σn kn δm,n = km. These expansions are explored in section (f) below for the two dual sets En, en and Un, un. (c) Covariant partner for En The covariant partner for En is given by (n)i = ij (En)j so that (n)i = ij (En)j = ij gja Rna = δi,a Rna = Rni Thus, one can regard the covariant vectors n as being the rows of matrix R R = = [1, 2, 3 .... N ]T which compare to S = [e1, e2, e3 .... eN ] (d) Summary of the basic facts: (en)k = Skn en em = 'nm |en| = = h'n S = [e1, e2, e3 .... eN ] (En)i = gia Rna En Em = g'nm |En| = R = [1, 2, 3 .... N ]T = g'na Sia en Em = δn,m En ≡ g'ni ei en = 'ni Ei (n)i = Rni e'n = R en where (e'n)i = δn,i // from Section 3 (b) In general, neither set of base vectors -- the tangent {..en .. } or the reciprocal {..En .. } -- is orthogonal, since the metric tensor g' in general is not diagonal. And in general none of these vectors is a unit vector. (e) Repeat the above for the inverse transformation: definition of the U'n Section 3 (d) introduced the inverse tangent base vectors called u'n. The prime indicates that these vectors exist in x'-space. In analogy with what was done above, one can define the inverse reciprocal base vectors U'n according to U'n ≡ gni u'i => u'n = ni U'i // since = g-1 Everything goes along as in the previous sections, but with these changes: g'↔ g R ↔ S en → u'n e'n → un En → U'n E'n → Un Here are the key results, translated from above, U'n ≡ gni u'i (SU'n) ≡ gni (Su'i) => Un = gni ui S = R-1 (U'n)i = gnaRia = g'iaSna // sum on second indices (un)k = Ski(u'n)i = δn,k (Un)k = Ski(U'n)i = Ski gnaRia = SkiRiagna = δk,agna = gnk (n)i = 'ij (U'n)j (u'n)k = Rkn u'n u'm = nm |u'n| = = hn R = [u'1, u'2, u'3 .... u'N ] (U'n)i = g'ia Sna U'n U'm = gnm | U'n| = S = [1, 2, 3 .... N ]T = gna Ria u'n U'm = δn,m U'n ≡ gni u'i u'n = ni U'n ('n)i = Sni un um = nm Un um = δn,m Un Um = gnm un = S u'n where (un)i = δn,i // from Section 3 (translated) It is helpful to keep all these eight vector symbol names in mind (and each has a covariant partner) x'-space x-space axis-aligned basis vectors e'n un (e'n)i = δn,i (un)i = δn,i dual partners to the above E'n Un (E'n)i = g'ni (Un)i = gni tangent base vectors u'n en (u'n)i = Rin (en)i = Sin reciprocal base vectors U'n En (U'n)i = g'ia Sna (En)i = gia Rna = gnaRia = g'naSia and recall that An am = δn,m for each of the four dual pairs (two primed, two unprimed). (f) Expanding vectors on different sets of basis vectors x-space expansions on un and Un Assume that V is some generic N-tuple V = (V1,V2....VN). There are various ways to expand V onto basis vectors. One way is to expand on the axis-aligned basis vectors un, which recall live in x-space, V = V1 u1 + V2 u2 +... = ΣnVnun where Un V = Vn Un = gni ui The components Vn are Un V because Un um = δn,m. From Section 5 (g), one can write Vn = gnmm ( regarded here as a definition of the m) so one finds that V = ΣnVn un = Σn gnm m un = Σnm gmn un = Σnm Um and thus another expansion for V is this V = 1 U1 + 2 U2 +... = ΣnnUn where un V = n Comments: 1. If V is not a contravariant vector, one can still define n = nmVm, but n won't be a covariant vector. A familiar example is that xn is never a contravariant vector if F is non-linear, but we can still talk about the components n ≡ nmxm . In Standard Notation, xn → xn and n → xn and we do not hesitate to use these two objects even though they are not tensorial vectors. 2. If V is a contravariant vector, the expansion above V = ΣnVnun displays the contravariant components of V. The second expansion V = Σnm Um is still an expansion for contravariant vector V, but it displays the components of the covariant vector which is the "partner" to V by n = nmVm. It would be incorrect to write this second expansion as = Σnm Um since that would say V = Σnm Um which is just not true. We comment later on how this situation changes a bit in the Standard Notation. x-space expansions on en and En Another possibility is to expand V on the tangent basis vectors en, and we denote the components just momentarily as αn, V = α1 e1 + α21 e2 +... = Σn αn en Using en Em = δn,m one finds that αn = En V = (En)k Vk = Rnk Vk = V'n // = 1 so AB = abAaBb = AkBk Therefore, the expansion is V = V'1e1 + V'2e2 +... = Σn V'n en where En V = V'n If it happens that the N-tuple V = (V1,V2....VN) transforms as a contravariant vector, then Vn are the contravariant components of that vector, and V'n are the contravariant components of V' in x'-space. On the other hand, if V is not a tensorial vector, so Vn are not components of a contravariant vector, we can still define V'n ≡ Rnk Vk, but then the V'n are not the contravariant components of V'. Writing V'n = g'nm'm the above expansion can be expressed as V = Σn V'n en = Σn,m g'nm'm en = Σm 'm Σng'mn en = Σm 'm Em so that V = '1E1 + '2E2 +... = Σn 'n En where en V = 'n Summary of x-space expansions: V = V1 u1 + V2 u2 +... = ΣnVn un where Un V = Vn Un = gni ui V = 1 U1 + 2 U2 +... = Σnn Un where un V = n V = V'1 e1 + V'2 e2 +... = Σn V'n en where En V = V'n En = g'ni ei V = '1 E1 + '2 E2 +... = Σn 'n En where en V = 'n Expanding on unit vectors. The covariant lengths of the different basis vectors are given by |en| = |e'n| = |un| =|u'n| = |En| = |E'n| = |Un| =|U'n| = Using these lengths, one can define unit vector versions of all the basis vectors and then rewrite the above expansions as expansions on the unit vectors with lower case coefficients. For example, using n ≡ en/ the third expansion above becomes (script font for unit-vector components) V = V'11 + V'22 +... = Σn V'n n where En V = V'n = V'n An example of a case where this last expansion would be useful is in the use of spherical curvilinear coordinates, where for example 1 = . The N-tuple (V'1, V'2 ... V'N), although related to contravariant vector V (V'n = Rnk Vk) , is not itself a contravariant vector since it does not obey the rule V'n = Rnk Vk . In fact V'n = Rnk Vk => (1/) V'n = Rnk (1/) Vk => V'n = Rnk (/)Vk x'-space expansions Having done x-space expansions, we turn now to x'-space expansions. These can be obtained from the x-space expansions by this set of rules, g'↔ g R ↔ S u'n → en un → e'n U'n → En Un → E'n V'n ↔ Vn 'n ↔ n and here then are the x'-space expansions: V' = V'1 e'1 + V'2 e'2 +... = ΣnV'n e'n where E'n V' = V'n E'n = g'ni e'i V' = '1 E'1 + '2 E'2 +... = Σn'n E'n where e'n V' = 'm V' = V1 u'1 + V2 u'2 +... = Σn Vn u'n where U'n V' = Vn U'n = gni u'i V' = 1 U'1 + 2 U'2 +... = Σn n U'n where u'n V' = n (g) Another way to write the En The reciprocal base vectors were defined above as linear combinations of the tangent base vectors, all in the general Picture A context, Ek ≡ g'ki ei It is rather remarkable that there is another way to write Ek in terms of the ei that looks completely different. In this other way, it turns out that Ek is expressed in terms of all the ei except ek and is given by (only valid in Picture B where g=1) Ek = det(R) (-1)k-1 e1 x e2 x ......x eN // ek missing k = 1,2,3...N This is a generalized cross product (Appendix A) of N-1 vectors, since ek is missing, so there are N-2 "crosses". The above multi-cross-product equation is a shorthand for (Ek)α ≡ det(R) (-1)k-1εαabc...x (e1)a(e2)b ...... (eN)x // κ and (eκ)K are missing Here ε is the totally antisymmetric tensor with N indices. If κ is the kth letter of the alphabet (k = 2 => κ = b ), then κ is missing from the list of summation indices of ε, and the factor (ek)κ is missing from the product of factors, so there are then N-1 factors. This cross product expression for En is derived in Appendix A. This is all fairly obscure sounding, but can be brought down to earth by writing things out for N = 3, where the formula reduces to this cyclic set of equations, E1 = det(R) e2 x e3 E2 = det(R) e3 x e1 E3 = det(R) e1 x e2 These equations can be verified in a simple manner. To show an equation is true, if suffices to show that the projections of both sides on the three en are the same, since the en form a complete basis as noted earlier. For the first equation one needs then to show that E1 en = det(R) e2 x e3 en for n = 1,2,3 If n=2 or n=3, both sides vanish, according to en Em = δn,m on the left, and according to geometry on the right, which leaves just the case n = 1. In this case the LHS = 1, so one has to show that e2 x e3 e1 = 1/det(R) = det(S) . But e1 e2 x e3 = (e1)i (e2 x e3)i = (e1)i εijk (e2)j(e3)k = εijk (e1)i(e2)j(e3)k = εijk Si1Sj2Sk3 = det(S) QED. The other two equations of the set can be verified in the same manner. Here is a picture (N=3) drawn in x-space for a non-orthogonal coordinate system. The vectors shown here form a distorted right handed coordinate system which has det(R) > 0. The reader is invited to exercise his or her right hand to confirm the directions of the arrows, E1 = det(R) e2 x e3 E2 = det(R) e3 x e1 E3 = det(R) e1 x e2 (h) Comparison of n and En One could create a covariant partner to en which would be n = en as described in Section 5 (g). This n is not the same as the reciprocal base vector En ≡ g'nk ek. The comparison is interesting: (n)i ≡ ik(en)k // matrix acts on vector index (En)i ≡ g'nk (ek)i // matrix acts on ek label If x-space is Cartesian, then n = en as usual, but of course En ≠ en in this case since g' ≠1. We mention this to head off a possible confusion when the Standard Notation is introduced in the next Section and the above two equations become (en)i ≡ gik(en)k // Standard Notation, g lowers an index (en)i ≡ g'nk (ek)i // Standard Notation, k is a label on ek, not an index The mapping to standard notation does not include n → en, for example. One fact about the standard notation is that, unlike the developmental notation, one cannot look at a vector in bold like en and determine whether it is contravariant or covariant. Only when the index is displayed can one tell. One can think of en as representing both its contravariant self and its covariant partner (en is a tensorial vector). (i) Handedness of coordinate systems: the en , the sign of det(S), and Parity Handedness of a Coordinate System. Let bn be a complete set of basis vectors in an N dimensional vector space, where the bn are not necessarily of unit length, and are not necessarily orthogonal. Consider this quantity B ≡ det (b1, b2.....bn ) = εabc..x (b1)a (b2)b.... (bN)x = b1 [b2 x b3......x bN] where the generalized cross product is discussed in Appendix A. This basis bn defines a "coordinate system" in that we can expand a position vector as follows x = Σn x(b)n bn where the x(b)n are the "coordinates" of point x in this coordinate system. We make the following definition: system bn is a "right handed coordinate system" iff B > 0 system bn is a "left handed coordinate system" iff B < 0 One of course wants to show that for N = 3 this definition corresponds to one's intuition about right and left handed systems. For N= 3 , B ≡ det (b1, b2, b3 ) = εabc (b1)a (b2)b(b3)c = b1 [b2 x b3] Suppose the bn are arranged as shown in this picture, where the visual intention is that the corner nearest the label b1 is closest to the viewer. With one's high-school-trained right hand, one can see that b2 x b3 points in the general direction of b1 (certainly b2 x b3 lies somewhere in the half space of the b2, b3 face plane which contains b1), and so the quantity B = b1 [b2 x b3] > 0. So this is an example of a right-handed coordinate system. The figure shown is a skewed 3-piped which can be regarded as a distortion of an orthogonal 3-piped for which the bn would span an orthogonal coordinate system in which one would have 1 = 2 x 3. The x'-space e'n coordinate system is always right handed. In our standard picture of x-space and x'-space, the coordinate system in x'-space is spanned by a set of basis vectors e'n where (e'n)i = δn,i , as discussed in Section 3 (a). This system is "right handed" because B = εabc..x (e'1)a (e'2)b.... (e'N) = εabc..x δ1aδ2b....δxN = ε123...N = +1 > 0 where we use the standard normalization of the ε tensor as shown. Notice that this conclusion is independent of the metric tensor g' in x-space. The x-space un coordinate system is always right handed. The basis un where (un)i = δn,i in x-space is right handed for the same reason as shown in the above paragraph, and for any g. When g=1 in x-space, the un form the usual Cartesian right-handed orthonormal basis in x-space. See Section 3 (c) concerning the meaning of "unit vector". The x-space coordinate en system handedness is determined by the sign of det(S). Our x-space coordinate system of great interest is that spanned by the en basis vectors, where (en)i = Sin as discussed in Section 3 (a). One has B = εabc..x (e1)a (e2)b.... (eN) = εabc..x Sa1 Sb2.... SxN = det(S) Therefore, using σ ≡ sign ( detS ) and J being the Jacobian as in Section 5 (k), system en is a "right handed coordinate system" iff det(S) = J > 0 or σ = +1 system en is a "left handed coordinate system" iff det(S) = J < 0 or σ = -1 The Parity Transformation. The identity transformation F = 1 results in matrix SI = I with detSI = +1. In this case the en form a right-handed coordinate system, and in fact en = un. The parity transformation F = -1, on the other hand, results in SP = -I with det(SP) = (-1)N and en = -un. When N is odd, the parity transformation converts the right-handed un system to a left-handed en system. If S is some matrix which does not change handedness, meaning detS > 0, then S' = SSP does change handedness for odd N, since in this case detS' = detS det SP = (-1)N detS. So given some S' that changes handedness for N=odd, one can regard it as "containing the parity transformation" which, if removed, would result in no handedness change. N=3 Parity Inversion Example. Since x' = -x under the parity transform F = -1, parity is a reflection of all position vectors through the origin. Objects sitting in x-space, such as N-pipeds, whether or not the origin lies inside the object, are "turned inside out" by the parity transformation, but the inside of the object still maps to the inside of the parity transformed object under this transformation. Consider this crude picture which shows on the left a right-handed 3-piped in x-space where e1 and e2 happen to be perpendicular, and the back part of the 3-piped is not drawn. This 3-piped is associated with some transformation S [ (en)i = Sin ] with detS > 0. Now consider S' = SSP = SPS = -S. For this S', the 3-piped appears as shown on the right. The two pipeds here are related by a parity transformation, all points inverting through the origin. On the right, the volume of the 3-piped lies toward the viewer from the plane shown. The circled dot on the left represents the out-facing normal vector of the 3-piped face area which is facing the viewer, and this normal is in the direction – e1xe2. This is called a "near face" in Appendix B since it touches the tails of the en. After the parity transformation, this same face has become the back face on the inverted 3-piped shown on the right, with out-facing normal indicated by the X. The direction of this normal is + e1xe2 . In general, an out-facing "near face" area points in the -En direction, and Appendix A shows that E3 = e1 x e2 / det(S). On the left we have E3 = e1 x e2 / |det(S)| so the face there just mentioned points in the -E3 = – e1xe2 direction. On the right we have E3 = e1 x e2 / det(S') = - e1 x e2 /|det(S)|, so the face there points in the -E3 = +e1xe2 direction. The sign of det(S) in the curvilinear coordinates application. For a given ordering of the x'i coordinates, det(S) will have a certain sign. By changing the x'i ordering to any odd permutation of the original ordering (for example, swap two coordinates), det(S) will negate because two columns of Sik(x') ≡ (∂xi/∂x'k) will be swapped. In the curvilinear coordinates application it is therefore always possible to select the ordering of the x'i coordinates to cause det(S) to be positive. One always starts with a right-handed Cartesian system for x-space, and det(S)>0 then guarantees that the en will form a right-handed system there as well. Since the underlying transformation F is assumed invertible, one cannot have det(S)=0 anywhere in the domain x (or range x') of x' = F(x), and therefore det(S) cannot change sign anywhere in the space of interest. For graphical reasons, we have selected coordinates in the "wrong order" in both the polar coordinates examples (called Example 1) and in the elliptic polar system studied in Appendix C, which is why detS < 0 for both these systems. 7. Translation to the Standard Notation In this Section we discuss the "translation" from our developmental notation (all lower indices; overbars for covariant objects) to the Standard Notation used in tensor analysis. The developmental notation has served well in the discussion of scalars and vectors, tensors of rank-0 and rank-1. For pure (unmixed) tensors of rank-2 it does especially well, allowing the use of matrix algebra to leverage the use of familiar matrix theorems such as det(ABC) = det(A)det(B)det(C) and A-1 = cof(AT)/det(A). The transformation of the contravariant metric tensor is cleanly expressed as g' = R g RT, and so on. The notation in fact works fine for unmixed tensors of any rank, but runs into big trouble with "mixed" tensors as shown in the next sections. (a) Outer Products It is possible to form larger tensors from smaller ones using the "outer product" method. For example, consider, Tab ≡ UaVb where U and V are assumed to be contravariant vectors. One then has T'ab = U'aV'b = (Raa'Ua') (Rbb'Vb') = Raa' Rbb' Ua'Vb' = Raa' Rbb' Ta'b' so in this way a contravariant rank-2 tensor (Section 5 (e)) has been successfully constructed from two contravariant vectors. Similarly, ab ≡ ab => 'ab = STaa' STbb' a'b' so the outer product of two covariant vectors transforms as a covariant rank-2 tensor. (b) Mixed Tensors and Notation Issues Suppose we take the "outer product" of a contravariant vector with a covariant vector, [ ... ]ab ≡ Uab where we are not sure what to call this thing, so we just call it [...]. Here is how this new object transforms (always: with respect to the underlying transformation x' = F(x) ) [ ... ]'ab = U'a'b = (Raa'Ua') (STbb'b') = Raa' STbb' Ua'b' = Raa' STbb' [...]ab This object transforms as a contravariant vector on the first index (ignoring the second), and as a covariant vector on the second index (ignoring the first). This is an example of a "mixed" rank-2 tensor. Extending this outer product idea, one can make elaborate tensor objects with an arbitrary mixture of "contravariant indices" and "covariant indices". For example [.....]abcd = Uab Xcd To write down the transformation rule for such an object, one must know which indices are contravariant and which are covariant. It is totally clear how the object transforms, looking at the right hand side of the equation, but somehow this information has to be embedded in the notation [.....]abcd because once this object is defined, the right hand side might not be immediately available for inspection. Worse, there may be no right hand side for a mixed tensor, because not all mixed tensors are outer products of vectors (they just transform as if they were). Just as we can use the idea ≡ V to convert a contravariant vector to its covariant partner, we can similarly use to convert the 1st or 3rd index on [.....]abcd from contravariant to covariant. We could apply two 's with the proper linkage of indices to convert them both at once. So given the ability of to change any index one way, and g to change it the other way, one can think of the 4-index object [.....]abcd as a family of 16 different 4-index objects, each corresponding to a certain choice for the indices being one type or the other. We know how to interconvert between these 16 objects just applying g or factors. So how does one annotate which of the 16 objects [.....] one is staring at for some choice of index types? Here is a somewhat facetious possibility, the Morse Code method ab ≡ Uab abcd = Uab Xcd Instead of having a bar over the entire object, in the first case the bar it is placed just over the right side of the W to indicate that b is a covariant index, while no bar means the first index is contravariant. The second example shows how horrible such a notation would be. We really want to put some kind of notation on the individual indices, not on the object! Here is a notation that is slightly better than the Morse code option, though similar to it, Wac = UaV XcY Here overbars on covariant indices distinguish them. Now one can dispense with the overbars on covariant vectors as well, putting the overbar on the index, for example a = abVb → V = g Vb . There are several problems with this scheme. One is that in the spinor application of tensor analysis used in special relativity ( see Section 5 (m) ), dots are placed on certain indices and these would conflict with the proposed overbars. A more substantial reason is that this last notation is hard to type (or typeset, as one used to say), it looks cluttered with all the overbars, and the subscripts are already hard to read without extra decorations since they are in a smaller font than the main text. (c) The up/down bell goes off This is where a bell went off somewhere, perhaps in the mind of Gregorio Ricci in the 1880-1900 time frame (1900 snippet quoted in section (j) below). Someone might have said: suppose, instead of using overbars on indices or some other decoration, we distinguish covariant indices by making them be superscripts instead of subscripts. Superscripts are as easy to type as subscripts, and the result is fairly easy to read and totally unambiguous. We would then have for our ongoing example, Wabcd = UaVbXcYd // a path not taken This is almost what happened, but the up/down decision went the other way and we now have: superscripts = contravariant = up subscript = covariant = down and then we get this translation Wac = UaV XcY → Wabcd = UaVbXcYd // the path taken and this has become The Standard Notation. Perhaps the reason for this choice was that the covariant gradient ∂n object appeared more commonly in equations than idealized objects such as dx, and ∂n already used a lower index. A downside of this particular up/down decision is that every student has be be confused by the fact that his or her familiar position, velocity and momentum vectors that always had subscripts suddenly have superscripts in the Standard Notation. The silver lining is that this shocking change alerts the student to the fact that whatever subject is being studied is going to have two kinds of vectors. Despite appearances, it is not completely obvious how one should translate the whole world as presented in the previous six Sections into this new notation. There are some subtle details that will be discussed in the following sections. (d) Some Preliminary Translations: raising and lowering indices on a vector with g In the entire rest of this entire Section, anything to the left of a → arrow is in "developmental notation", while anything to the right of → is in "Standard Notation". So we start translating some of the results above: s → s // a scalar Va → Va // a contravariant rank-1 tensor (vector) a → Va // a covariant rank-1 tensor (vector) Mab → Mab // a contravariant rank-2 tensor ab → Mab // a covariant rank-2 tensor gab → gab // the contravariant rank-2 metric tensor ab → gab // the covariant rank-2 metric tensor g is inverse of → gab is inverse of gab As noted earlier, one "feature" of the Standard Notation is that it is no longer sufficient to specify an object by a single letter. One has to somehow indicate the index nature by showing index positions. Thus, "g" stands for all four metric tensors gab , gab, gab and gab. The pure covariant metric tensor is gab or perhaps g** . At first this seems a disadvantage of the notation, but one then realizes that the true object really is "g", and it has four different "representations" and the notation makes this very clear. Still, one cannot just write det(g) because det(g) is representation dependent, so one must say something like det(gab) or det(g**) to denote a particular determinant. As for converting a vector from one type to the other, a = abVb → Va = gabVb // gab "lowers" a contravariant index Va = gab b → Va = gab Vb // gab "raises" a covariant index , and so in this new notation, the covariant metric tensor gab becomes an "index lowering operator" and the contravariant metric tensor gab becomes an "index raising operator". This is a huge advantage of the Standard Notation. It pretty much eliminates the need to think, something universally appreciated. In a certain obscure sense, it is like double entry accounting (credits and debits), where the notation itself serves as a check on the accuracy of bookkeeping entries, as will be seen below. As for bolded vectors, the translation rule is, V → V → V The reason is that the overbar is no longer used to denote covariancy. The above lines show a subtle change in the interpretation of the bolded symbol V in the standard notation: the single symbol V stands for both the developmental vector V and for its developmental covariant partner vector . The new symbol V is both contravariant with components Vn and it is covariant with components Vn. The invariant distance and covariant dot products: dxi → dxi (ds)2 = ab dxa dxb → gabdxadxb AB = ab Aa Bb → AB = gab Aa Bb = AbBb = AaBa = gabAaBb The general idea is this: any tensor index on any tensor object can be raised by gab and can be lowered by gab. Remember that a tensor object lives in some space like x-space, so we shall have to ponder what to do for our matrices Sab and Rab which live half in x-space and half in x'-space, a subject we defer for a short while. (e) Contraction of a Pair of Indices When two indices are summed together in a tensor expression and one is up and the other down, one says that the two indices are contracted. Here is an example, where the index b is contracted, Va = gabVb It is shown below that contracted indices neutralize each other in terms of how an object transforms. Thus, for example, the RHS above gabVb transforms as a covariant vector, which conveniently matches the LHS. Similarly, AaBa transforms as a scalar. (f) Dealing with the matrix R Consider the translation of this partial derivative into the new up/down notation. Since the differential dx element is contravariant and is now written dxi , (∂x'i/∂xk) → (∂x'i/∂xk) In terms of "existence", this object has one leg in each space of Picture A. The gradient operator ∂/∂xk is an x-space thing, while x'i is an x'-space thing. Since this object does not live in x-space or in x'-space exclusively, but straddles the two spaces, it cannot possibly be a tensor of any kind. Recall that a tensor object must be entirely within a space, it cannot have body parts hanging out into other spaces. Nevertheless, it seems clear that each of the two indices has a well-defined nature. We showed that the gradient is a covariant vector, so we regard k as a covariant index. And of course dx'i is a contravariant vector, so i is a contravariant index. Here then is the proper translation: Rik ≡ (∂x'i/∂xk) → Rik ≡ (∂x'i/∂xk) To summarize, Rik is not a mixed rank-2 tensor, though it looks just like one. Therefore , Rik can never appear in a tensor equation -- it just appears in the equations that show how tensors transform. However, each of the two indices of R has a well-defined transformational nature, and we place them up and down in the proper manner. It is very typical for an object to have up and down indices but the object is not a tensor. The canonical example is that for a non-linear transformation x' = F(x), xi has a contravariant index but is not a contravariant vector. Consider now the translation of the transformation rule for a contravariant vector V'a = RabVb → V'a = RabVb // contravariant Even though R is not a tensor, we see that index b is contracted and is thus neutralized from the evaluation of the tensor nature of the RHS. This leaves upper index a as the only free index, indicating that the RHS is a contravariant vector, and this of course then matches the LHS. So we can deal with the indices on R just as we deal with indices on true tensors. Notice that, even though both sides of V'a = RabVb have the same "tensor nature" (both sides are a contravariant vector) one cannot ask how the equation V'a = RabVb "transforms" under a transformation. That question can only be asked about equations constructed of objects all of which are tensors in the same space. Here V and half of R are in one space, and V' and the other half of R are in a different space. There is no object called R', as if R were in x-space and R' were in x'-space. (g) Repeat the above section for S We omit the words and just show the translations (∂xi/∂x'k) → (∂xi/∂x'k) Sik ≡ (∂xi/∂x'k) → Sik ≡ (∂xi/∂x'k) 'a = STabb = Sba b → V'a = SbaVb // covariant (h) About ε and δ The Kronecker δ is sometimes written in different ways to make things "look nice", δab = δab = δba = δba = δa,b Section (m) will show that one can regard the above sequence of equalities as saying gab = gab = gba = gba = δa,b where these g objects are mixed versions of the symmetric rank-2 metric tensor gab. There is no "δ tensor", it is the g tensor, but tradition is to write the diagonal objects using the δ symbol. The object εabc... is a bit more complicated. It can at first be regarded as a mere bookkeeping device, in which context it is usually called "the permutation tensor". It appears for example in the expansion of a determinant det(M) = εabc...xM1aM2b.....MNx = εabc...xMa1Mb2.....MxN or in an ordinary cross product Aa = εabcBbCc . This permutation tensor has the usual properties that ε123...N = +1 , that ε changes sign when any two indices are swapped, and that ε vanishes if two or more indices are the same. This permutation "tensor" is not really a tensor since one would regard it as being the same in x-space or x'-space. Whether indices are written up or down on this ε is immaterial. At another level, however, εabc...x with N indices (the same ε symbol is used) is a covariant rank-N tensor density of weight -1 known as the Levi-Civita tensor. This subject is addressed in Appendix D in much detail. In what we call the Weinberg convention, individual indices of ε can be raised and lowered by g as discussed in section (d) just as with any tensor. Therefore, in Cartesian space with g = 1, indices on ε are raised and lowered with no consequence, and then one can identify any form of ε as being the permutation tensor. For example, εabc= εabc = εabc and so on. In a non-Cartesian x-space, however, one would say that εabc = gbb'εab'c ≠ εabc. In the Weinberg convention, one sets ε123..N = ε'123..N = 1 and εabc..x = ε'abc..x has the properties of the permutation tensor described above and these properties are the same in x-space as in x'-space. Then for general g≠1, εabc..x is NOT the permutation tensor. The bottom line is that one must be aware of the space in which one is working (the Picture). The ε appearing above in the determinant expansion is always just the permutation tensor, but in the cross product that is not the case, and one would properly write Aa = εabcBbCc and conclude that the cross product of two ordinary contravariant vectors is a covariant vector density (Appendix D (g)). Again, in Cartesian space where one often works, this would be the same as Aa = εabcBbCc = εabcBbCc , but the "properly tilted form" Aa = εabcBbCc reveals the tensor nature of the object Aa. As mentioned below in section (u), this "covariant" equation would appear as A'a = ε'abcB'bC'c in x'-space, but since A'a is a covariant vector density, A'a ≠ RabAb, and in fact A'a = |J| RabAb . The permutation tensor εabc... and the contravariant Levi-Civita tensor εabc...x are both "totally antisymmetric" which just means ε changes sign if any pair of indices is swapped. In fact, as discussed in Appendix D (c), there IS only one antisymmetric tensor of rank N apart from a multiplicative scalar factor, and εabc...x is it. This fact simplifies various calculations. Technically, εabc...x is a totally antisymmetric tensor density, but normally it is just called "the totally antisymmetric tensor". As shown in Appendix D, the covariant Levi-Civita tensor εabc...x is also totally antisymmetric and is therefore a multiple of εabc...x. The reader is invited to peruse Appendix D at some appropriate time for more about tensor densities and the ε tensor. (i) Further development of the Standard Notation This long section contains a veritable grab-bag of important Standard Notation facts. Each such fact is proven and not just quoted. Many of the results presented here anticipate more formal presentations of the same results in later sections. Translation of determinants. Section (g) showed that Rab → Rab, so one translates from old to new notation, det(R) = εabc...R1aR2b....RNx → det(Rij) = εabc...R1aR2b....RNx det(S) = εabc...S1aS2b.....SNx → det(Sij) = εabc...S1aS2b....SNx det(R) = εabc...Ra1Rb2....RxN → det(Rij) = εabc...Ra1Rb2....RxN det(S) = εabc...Sa1Sb2.....SxN → det(Sij) = εabc...Sa1Sb2....SxN where ε is the bookkeeping permutation tensor discussed in section (h). Inverse of R and S. Again, section (g) showed that Rab → Rab. In the standard notation, imagine that there is some inverse R-1 defined by (R-1)caRab = δcb. The chain rule says that (∂xc/∂x'a) (∂x'a/∂xb) = δcb or Sca Rab = δcb and therefore it must be that (R-1)ca = Sca. A similar argument shows that (S-1)ca = Rca. Using the results of the next section which allow us to raise and lower indices on both sides of an equation, this relationships R-1 = S is valid for all four matrix position possibilities, (R-1)ik = Sik (R-1)ik = Sik (R-1)ik = Sik (R-1)ik = Sik and of course the same is true for S-1 = R. Thus arise these translations from old to new notation: R-1 = S → (R-1)ik = Sik and all other index combinations S-1 = R → (S-1)ik = Rik and all other index combinations RR-1 = RS = 1 etc → Rik(R-1)ka = RikSka = δia etc (1) Tensor g raises and lowers any index. So far the following translation rules have been established: gab → gab ab → gab Rik → Rik Sik → Sik It was shown in developmental notation Section 5 (e) how a rank-2 contravariant tensor transforms. Here then is how that statement translates to the new notation M'ab = Raa'Rbb'Ma'b' → M'ab = Raa'Rbb'Ma'b' contravariant rank-2 tensor (2) 'ab = Sa'aSb'ba'b' → M'ab = Sa'aSb'bMa'b' covariant rank-2 tensor Since g is such a rank-2 tensor, replace M by g to get g'ab = Raa'Rbb'ga'b' → g'ab = Raa'Rbb'ga'b' (3) 'ab = Raa'Rbb'a'b' → g'ab = Sa'aSb'b ga'b' It was shown in section (d) that Va = gaa'Va' and Va = gaa' Va' so that gaa' lowers a vector index and gaa' raises a vector index. That is to say, gaa' converts a contravariant vector index into a covariant one, and gaa' does the reverse. What does gaa' do to the index of a rank-2 tensor? Consider the following definition: Mab ≡ gaa'Ma'b (4) Since gab and gab are inverses, it follows that Mab = gaa' Ma'b (5) How does this new object Mab transform? The claim is that it transforms as a mixed rank-2 tensor, which would mean that M'ab = Raa' Sb'b Ma'b' The upper index gets a factor Raa' and the lower index gets a factor Sb'b , consistent with (*) above. It is not hard to prove this claim: (4) (3) (2) (5) M'ab = g'aa'M'a'b = ( RacRa'd gcd ) ( Sea'SfbMef) = ( RacRa'd gcd ) ( Sea'Sfb gei Mif ) = [RacRa'd gcd Sea'Sfb gei] Mif = [Rac (Sea'Ra'd) gcd Sfb gei] Mif = [Rac (SR)ed gcd Sfb gei] Mif = [Rac δed gcd Sfb gei] Mif = [Rac gcd Sfb gdi] Mif (1) = [Rac (gcd gdi) Sfb] Mif = [Rac δci Sfb] Mif = [Rai Sfb] Mif inverses = Raa' Sb'b Ma'b' QED Similarly one could define Mab ≡ gaa'Ma'b and one would find that M'ab = Sa'a Rbb' Ma'b' so Mab is then another member of the family of rank-2 tensors. Finally were one to define Mab ≡ gbb'Mab' one would find that Mab transforms as in (2). To summarize the four transformation results M'ab = Raa' Rbb' Ma'b' M'ab = Raa' Sb'b Ma'b' M'ab = Sa'a Rbb' Ma'b' M'ab = Sa'a Sb'b Ma'b' One sees then a family of four tensors associated with M. One is contravariant, one is covariant, and the other two are mixed. Let [----i---] represent a tensor with a certain contravariant index i and dashes indicate other indices which could be up or down. Similarly define [----i---] as another tensor in the same family where the index i that was up is now down. Raising Lowering Rule: gaa' [----a'---] = [----a---] and gaa' [----a'---] = [----a---] The notion of higher rank tensors is coming soon, but we just want to establish the general idea that ANY index on ANY tensor can be raised or lowered by an appropriate g tensor. For the rank-2 tensors this was demonstrated explicitly above, and section (d) showed it was valid for rank-1 tensors (vectors), gaa'Va' = Va gaa' Va' = Va Comment: Notice that in every equation shown above, the summed indices always occur in the contracted form discussed in section (e) above, which is to say, one index is up and the other is down. Contraction Tilt Reversal Rule: [-----a---------a----] = [-----a---------a----] This is proved in section (k) below, but since we are going to need it right now, here is a preview of that proof: ( note that gab gac = gba gac = δbc ) [-----a---------a----] = gab gac [-----b---------c----] = δbc [-----b---------c----] = [-----b---------b----] The upshot is that one can "reverse the tilt" on any pair of contracted indices. The Diagonal g Rule: gab = δab and gab = δab // and same for g' This is proved in section (m) below, but since we are going to need it right now, here is a preview of that proof: gab = gaa' ga'b // gaa'raises the first index of tensor ga'b = δab // because gij and gij are inverses of each other Matrix Multiplication in the Standard Notation. Although there are various forms of matrix multiplication, the most standard form is that obtained when all matrices have a "down-tilt" form. Consider this example, Cab = AacBcb where it is assumed that all three objects are down-tilt rank-2 tensors (or objects like Rac and Sac whose indices behave as if they rank-2 tensors). Although these are "split level" matrices, once can see that the index c has the correct "adjacency" property to justify matrix multiplication. A second requirement is that any matrix summed index must be a genuine contraction with one index up and the other down. The above tensor transformation rule can then be written in this more compact matrix notation, C = AB // all down-tilt An example appears in (1) above: δia = RikSka = ↔ 1 = RS S-1 = R By application of suitable g tensors to the first equation above (or by lowering index a and raising index b), one gets Cab = AacBcb . Application of the Tilt Reversal Rule on index c then gives Cab = AacBcb Again the adjacency and contraction rules are met, so this equation can be represented also by C = AB // all up-tilt and so δia = RikSka = ↔ 1 = RS S-1 = R Comment: Notice that the equation 1 = RS is valid both in the standard notation (providing R,S and 1 all have the same tilt) and in the developmental notation. The upshot is that in Standard Notation, matrix notation can be used if all matrices in the equation being represented are either all down-tilt or all up-tilt. As will be shown later, such matrix equations are all "covariant" in that both sides of the equation have the same tensor transformation property, and this is due to the fact that the matrix summation index is a contraction. One could talk about matrix multiplication in other cases, such as Cab = AacBcb but since index c is not a contraction, if A and B are tensors, then C cannot be a tensor and we don't even want to think about such equations. Transpose of a rank-2 tensor. If A is a rank-2 tensor, the translation mapping (AT)ab = Aba → (AT)ab = Aba seems obvious, and the object AT therefore also transforms as a rank-2 tensor. Once (AT)ab = Aba is established in standard notation, one can apply the metric tensor g to lower either or both of the indices of this equation, to get (AT)ab = Aba , (AT)ab = Aba, and (AT)ab = Aba. Notice in all four equations that the indices on the two sides of the equation are reflected in a vertical axis passing between the indices. This causes a left index to become a right index and vice versa, as one would expect for transposing a matrix. Moreover, on each side of all four equations, each index has the same contravariant/covariant sense. This same argument also applies to R and S even though they are not tensors. The only difference is that the first index of Rba is lowered by g' while the second by g. For example, (RT)ab = Rba where b is a g' type index and a is a g type index, as will be elaborated in section (o) below. Similarly (ST)ab = Sba . To summarize the situation with transposes in Standard Notation: (AT)ab = Aba (RT)ab = Rba (ST)ab = Sba (AT)ab = Aba (RT)ab = Rba (ST)ab = Sba (AT)ab = Aba (RT)ab = Rba (ST)ab = Sba (AT)ab = Aba (RT)ab = Rba (ST)ab = Sba Notice that the rule is not (AT)ab = Aba which would be a straight swap of indices (and would result in AT not being a tensor). The straight swap idea works for the pure contravariant and pure covariant forms of A, but not for the tilted forms! As shown in Theorem 4 below, this tilted transpose form has some interesting implications. Theorem 1: For either the down-tilt or up-tilt version of S, S is a real-orthogonal matrix. This means all of the following: S-1 = ST SST = 1 STS= 1 Proof of theorem: start with (3) above which says gij is a covariant rank-2 tensor, g'ab = Sa'aSb'b ga'b' // next, apply gb'b to both sides (or just raise index b on both sides) g'ab = Sa'aSb'b ga'b' // next, reverse the tilt of the b' index g'ab = Sa'aSb'b ga'b' // next, use the Diagonal g rule in two places δab = Sa'aSb'b δa'b' = Sa'aSa'b // next, use (ST)aa' = Sa'a δab = (ST)aa'Sa'b // next use matrix notation as per above 1 = STS => ST = S-1 => 1 = SST Theorem 2: For either the down-tilt or up-tilt version of R, R is a real-orthogonal matrix. This means all of the following: R-1 = RT RRT = 1 RTR= 1 In other words, the previous theorem also applies to R. Again, this fact is not true of the developmental notation matrix Rab. The proof is very similiar but just different enough to warrant showing it : Proof of theorem: start with (3) above which says gij is a covariant rank-2 tensor, g'ab = Raa'Rbb'ga'b' // next, apply gb'b to both sides (or just lower index b on both sides) g'ab = Raa'Rbb'ga'b' // next, reverse the tilt of the b' index g'ab = Raa'Rbb'ga'b' // next, use the Diagonal g rule in two places δab = Raa'Rbb'δa'b' = Raa'Rba' // next, use (RT)a'b = Rba' δab = Raa'(RT)a'b // next use matrix notation as per above 1 = RRT => RT = R-1 => 1 = RTR Comment: In the developmental notation, neither R nor S is a real-orthogonal matrix, unless by accident, but in the tilted standard notation they both are! Theorem 3: For either the down-tilt or up-tilt version of S, (a) S = RT and R = ST (b) Sab = Rba and Sab = Rba ( reflect indices in vertical line between them) Proof of theorem: From the inverse discussion at the start of this section, S = R-1 for any index positions in the standard notation, including the up-tilt and down-tilt positions. From Theorem 2, RT = R-1 in either up-tilt or down-tilt forms. Therefore S = RT (hence ST = RTT= R) in either up-tilt or down-tilt form. Thus S = RT => Sab = (RT)ab = Rba S = RT => Sab = (RT)ab = Rba QED An implication of this theorem is that one can completely eliminate references matrix S in tensor analysis and that is what is usually done. This is like replacing S with R-1 in the devepmental notation. Theorem 4: For a standard notation tilted matrix, det(A) ≠ det(AT). This surprising result points out a potential hazard of using matrices in the standard notation, and perhaps is an indication of why people avoid matrix concepts and just write out all the components. For a traditional matrix Aab the determinant is given by either of these forms (figure on rows or columns) det(A) = εab..A1aA2b.... = εab..Aa1Ab2 ... det(AT) = εab..AT1aAT2b.... = εab..Aa1Ab2...... = det(A) In the tilted standard notation, however, one has det[Aij] = εab..A1aA1b.... = εab..Aa1Ab2 ... det[(AT)ij] = εab.. (AT)1a(AT)1b.... = εab...Aa1Ab2 but this is not the same as either of the det[Aij] forms! Here is a simple example: det(A) = εab..A1aA1b... = = A11 A22 – A21 A22 det(AT) = εab...Aa1Ab2 = = A11 A22 – A21 A12 ≠ det(A) The point is that the standard notation transpose rule (AT)ij = Aij doesn't just swap the indices, it also changes the tilt. That means that the columns of one determinant matrix are not the same as the rows of the other and that is why det(A) ≠ det(AT). Corollary: RTR = 1 does not imply that det(R) = ± 1. The usual proof goes that det(RTR) = det(RT)det(R) = det(R)det(R) = [ det(R) ]2 = 1, but of course the part saying det(RT) = det(R) is no longer valid. Therefore, the fact that RTR = 1 does not lead us to the false conclusion that the Jacobian of Section 5 (k) is somehow forced to be ± 1 ! So although R and S are real orthogonal matrices, they are not "rotation matrices". Orthogonality rules. The first two rules are easy to show, RTR = 1 => (RT)abRbc = δac => RbaRbc = δac RRT = 1 => Rab(RT)bc = δac => RabRcb = δac Lowering a and raising c on both sides and reversing the b tilt then gives the other two rules RbaRbc = δac RabRcb = δac These orthogonality rules are derived in a slightly different manner and order in section (r) below. Variations on the relation between g and g'. Item (3) above gave the basic statement of the transformation properties of tensor g. These can be inverted as follows g'ab = Raa'Rbb'ga'b' => gab = (R-1)aa'(R-1)bb'g' a'b' = Saa' Sbb' g' a'b' g'ab = Sa'a Sb'b ga'b' => gab = (S-1)a'a (S-1)b'b g'a'b' = Ra'a Rb'b g'a'b' and here then is a summary, g'ab = Raa'Rbb' ga'b' gab = Saa'Sbb' g'a'b' g'ab = Sa'a Sb'b ga'b' gab = Ra'a Rb'b g'a'b' If x-space is Cartesian with g = 1, the first column above simplifies to g'ab = RacRbc // g = 1 g'ab = Sca Scb = Rac Rbc // g = 1 Notice that the summation index c is not a contraction here. Also, although R and S are not tensors, the sums shown produce the true tensors g'ab and g'ab. (j) Tensors of Rank n, direct products, Lie groups, symmetry and Ricci-Levi-Civita The most general tensor of rank n (aka order n) will have some number s of contravariant indices and then some number n-s of covariant indices. If s = n, the tensor is pure contravariant, and if s = 0, it is pure covariant, otherwise it is "mixed" (as opposed to "pure"). The transformation of the tensor under F will show a factor Raa' for each contravariant index, and a factor Sa'a ( = Raa' as shown below in section (q)) for each covariant index, as illustrated by this example : T'abcde = Raa' Rbb' Rcc' Sd'd Se'e Ta'b'c'd'e' T'abcde = Raa' Rbb' Rcc' Rdd' Ree' Ta'b'c'd'e' Note that for Raa' and Raa' the second index is the summation index, but for Sa'a it is the first index. A rank-n tensor always transforms the way an outer product of n vectors transforms if those vectors have indices which type-match those of the tensor. In the above case, an object that would transform the same as Tabcde would be AaBbCcDdEe A tensor of rank-n has 2n tensor objects in its family since each index can be up or down. For example, the tensor T above is one of 25 = 32 tensors one can form. Of these, one is pure covariant and one is pure contravariant and 30 are mixed. If any of these tensors is a tensor field, such as Tabcde(x), then of course all family members are tensor fields. Direct Products. Consider again the outer product of vectors AaBbCcDdEe. The transformation A'a = Raa'Aa' occurs in an N-dimensional contravariant vector space we shall call R . In this space one could establish a set of basis vectors, and of course there are rules for adding vectors and so on. Transformation B'a = RacBc occurs in an identical copy of the space R, but transformation D'd = Sd'dDd' = Rdd'Dd' occurs in a covariant version of R we call . Since dot products (inner products) have been established for vectors in these spaces, they can be regarded as full blown Hilbert Spaces with the caveats of Section 5 (i). The transformation of the outer product object, as already noted, is given by A'aB'bC'cD'dE'e = Raa' Rbb' Rcc' Rdd' Ree' Aa'Bb'Cc'Dd'Ee' and one can consider the operator Raa' Rbb' Rcc' Rdd' Ree' as a transformation element in a so-called direct product space which in this case would be written Rdp = R R R One could then define (Rdp)abcde ; a'b'c'd'e' ≡ Raa' Rbb' Rcc' Rdd' Ree' so that A'aB'bC'cD'dE'e = (Rdp)abcde ; a'b'c'd'e' Aa'Bb'Cc'Dd'Ee' and of course this would apply to any tensor of the same index configuration, such as T'abcde = Rdpabcde ; a'b'c'd'e' Ta'b'c'd'e' This suggests a definition of "tensor" as follows" : tensors are those objects that are transformed by all possible direct product representations formable from the two fundamental vector representations R and . To this set of spaces one would add the identity space 1 to handle tensorial scalars. Appendix E continues this direct product discussion in terms of the basis vectors that form a complete set for a direct product space such as Rdp and shows how to expand tensors on such bases. Lie Groups. The direct product notion is just a formalism, but the formalism has some implications when the space R is associated with a "representation" of a Lie group. In this case, a direct product Rdp = R R can be written as a sum of "irreducible" representations of that group. What this means is that the transformation elements of Rdp and the objects Tab can be shuffled around with linear combinations so that (Rdp)aba'b', when thought of as a matrix with columns labeled by N2 ab possibilities and rows labeled by the N2 a'b' possibilities, appears in "block diagonal form" with all zeros outside the blocks. In this case, the shuffled components of tensor Tab can be regarded as a non-interacting assembly of pieces each of which transforms according to one of those blocks of the shuffled (Rdp)aba'b'. The most famous example occurs with N=3 and the rotation group SU(2) in which case R(1) ≡ R can be decomposed according to R(1) R(1) = R(2) R(1) R(0) where the symbols indicate this block diagonal form. In this case the blocks are 5x5, 3x3 and 1x1, fitting onto the diagonal of the 9x9 matrix area. The numbers L = 0,1,2 here label the rotation group representations and that label is associated with angular momentum. The elements of the 5x5 block are called D(2)M,M'(φ,θ,ψ) where M,M' = 2,1,0,-1.-2, and where φ,θ,ψ are the "Euler angles" which serve to label a particular rotation. This D(2)object is the L=2 matrix representation of the rotation group. Taking two vectors A and B, one can identify AB as the combination transforming according to R(0) ("scalar") and AxB ( linearly combined) as that transforming as R(1) ("vector"). The traceless matrix AiBj - δi,jAB has 5 independent elements associated with R(2) ( "quadrupole"). This whole reduction idea can be applied to larger direct products such as R(1) R(1) R(1) and tensor components Tabc. The Standard Model of elementary particle physics is chock full of direct products of this nature, where the idea of rotational symmetry is extended to other kinds of "internal" symmetry, spin and isospin being two examples. Representations of the Lie symmetry group SU(3) are associated with quarks which are among of the fundamental building blocks of the Standard Model. The group discussion above can be applied generally to quantum physics. The basic idea is that if "the physics" (the Hamiltonian or Lagrangian) describing some quantum object is invariant under a certain symmetry group (such as rotational symmetry or perhaps some discrete crystal symmetry), then the quantum states of that object can be classified according to the representations of that group. The Bohr hydrogen atom "physics" H ~ 2-1/|r| has perfect rotation group symmetry and is also symmetric about the axis (angle ψ) from center to electron (no spin). The representation functions then must have M' = 0, and then D(L)M,0(φ,θ,ψ) ~ YLM (θ,φ), the famous spherical harmonics that describe the "orbitals" which have mystified first-year chemistry students for the last 100 years. Historical Note: Ricci and Levi-Civita (see Refs) referred to rank-n tensors as "systems of order n" and did not include mixed tensors in their 1900 paper. Nor did they use the Einstein summation convention, since Einstein thought of that later on. They did use the up and down index notation pretty much as it is used today, though the up indices are enclosed in parenthesis. Here is a direct quote from the paper where the nature of the contravariant and covariant tensors is described (with crude translation below for non-French readers). Equation (6) had typos which some thoughtful reader corrected: the y subscripts should be r's and the x subscripts should be s's. In our notation ∂xs/∂yr → ∂xs/∂x'r = Ssr = Rrs. We will say that a system of order m is covariant (and in this case we will designate its elements by the symbol Xr1,r2....) (r1,r2.... can each take all the values 1...n), if the elements Yr1,r2.... of the transformed system are given by the formulas (6) . We will designate on the contrary by the symbols X(r1,r2....) the elements of a contravariant system, which is to say of a system where the transformation is represented by the formulas (7) , the elements X and Y being related respectively to (presumably "are functions of") the variables x and y. Their y is our x', and their n is our N. Notice that the indices on the coordinates themselves are taken down, contrary to current usage. They do not explain why the words contravariant and covariant are used. (k) The Contraction Tilt-Reversal Rule In some complicated combination of multiple tensors and perhaps some R and S objects, imagine there is somewhere a pair of summed indices where one is up and the other is down. As noted above, such a sum is called a contraction. The contracted indices could be on the same object or they could be on different objects. We depict this situation with the following symbolic notation, [-----a---------a----] where the dashes indicate indices that we don't care about and which won't change -- each one could be up or down. We know we can reverse the tilt this way, [-----a---------a----] = gab gac [-----b---------c----] where the first g raises the index b to a, and the second g lowers the index c to a. But the two g's are inverses, gab gac = gba gac = δa,c, which at once gives the desired result [-----a---------a----] = [-----a---------a----] // the Contraction Tilt-Reversal Rule A notable example of course is this: AaYa = AaYa = " A Y " // or perhaps " A.Y " as noted in Section 5 (i) (l) The Contraction Neutralization Rule A contracted index pair plays no role in how an object transforms, the two indices neutralize each other, as we now show. First, recall that the indices on a general rank-n tensor (perhaps formed from several tensors) transform the same way an outer product of n vectors transforms, where the vector index types match those of the tensor. The vectors transform this way: V'a = RabVb V'a = SbaVb So, we take our same "big object" above and now ask how it transforms. In the following, the X's represent either R or S factors for the dash indices (each of which might be up or down): [-----a---------a----]' = XXXXX Rab XXXXXXXXX Sca XXXX [-----b---------c----] = Sca Rab XXXXX XXXXXXXXX XXXX [-----b---------c----] = δcb XXXXX XXXXXXXXX XXXX [-----b---------c----] = XXXXX XXXXXXXXX XXXX [-----a---------a----] where now the only X's left are for the other indices. Again we look at our canonical example, AaYa = AaYa = A.Y The contracted vector indices cancel each other out and the resulting object transforms as a scalar. Here are some examples of tensor transformations with 0,1 and 2 index pairs contracted: T'abcde = Raa' Rbb' Rcc' Sd'd Se'e Ta'b'c'd'e' // no pairs contracted T'abcae = Rbb' Rcc' Se'e Ta'b'c'a'e' // index a contracted T'abcab = Rcc' Ta'b'c'a'e' // index a and index b contracted Q' = Q where Q = AaYa and Q' = A'aY'a // index a contracted This shows the idea that one can take a larger tensor like Tabcde and form from it smaller (lower rank) tensors by contracting tilted pairs of indices. In the above example list we really have Dbce ≡ Tabcae = a mixed rank-3 tensor Ec = Tabcab = a contravariant vector (rank-1 tensor) Q = AaYa = a scalar (rank-0 tensor) It is similarly possible to build larger tensors from smaller ones, for example Zabcde = Va We gab Lc which goes under the same rubric "outer product" mentioned earlier. (m) Raising and lowering indices on g On the one hand, since gab and gab are inverses of each other (formerly and g) , one has gabgbc = δa,c = δac where the above-mentioned "look-nice" form of δa,c makes indices match. On the other hand, gabgbc = gbc // left g lowers the left index of the right g, or the converse Comparison shows that gbc = δbc As a sanity check, consider Va = gabVb Applying our Contraction Tilt-reversal Rule, this can be written Va = gabVb but this is = δabVb = Va There are many ways to write things, here is a collection (gab = gba !) gabgbc = δac = gac // start out gabgbc = δac = gac // tilt reversal of the line above; just says δabδbc = δac gabgbc = δac = gac gabgbc = δac = gac (n) Other forms of R Object Rab = (∂x'a/∂xb) was considered above. One could lower the a index using g'** since x'a is in x'-space and is an up index. The index in ∂/∂xb = ∂b is really a lower index (gradient), so one could in effect raise it using g** (no prime) because ∂/∂xb is in x-space. So when raising and lowering indices on Rab one has the unusual situation that one must use g' when acting on the first index, and g when acting on the second. With this in mind, we can now write three other index configurations of Rab Rab = ( ∂x'a/∂xb) // original object (formerly Rab) Rab = Rab' gb'b = (∂x'a/∂xb) // g pulls up the second index of Rab Rab = g'aa'Ra'b = (∂x'a/∂xb) // g' pulls down the first index of Rab Rab = g'aa'Ra'b' gb'b = (∂x'a/∂xb) // both actions at once Although the g and g' factors can be placed anywhere, we have put g' factors on the left of R, and g factors on the right, each next to its appropriate leg of R. In each case, examination of the corresponding partial derivative shows that that the index sense matches on both sides. For example, in Rab = (∂x'a/∂xb) = ∂bx'a, both indices are contravariant on both sides. Remember that Rab is not a contravariant rank-2 tensor due to its dual-space nature. (o) Summary of facts about R Rik ≡ (∂x'i/∂xk) → Rik ≡ (∂x'i/∂xk) V'a = RabVb → V'a = RabVb Va = SabV'b [= RbaV'b] Rab(x) = (∂x'a/∂xb) // original object (formerly Rab) Rab = Rab' gb'b = (∂x'a/∂xb) // g pulls up the second index Rab = g'aa'Ra'b = (∂x'a/∂xb) // g' pulls down the first index Rab = g'aa'Ra'b' gb'b = (∂x'a/∂xb) // both actions at once Rab = g'aa'Ra'b' gb'b // the inverse of the previous line (using gabgbc = δac twice) (p) Repeat all the above for S Sik ≡ (∂xi/∂x'k) → Sik ≡ (∂xi/∂x'k) 'a = STab b = Sba b → V'a = SbaVb Va = RbaV'b [= SabV'b] Sab(x) = (∂xa/∂x'b) // original object (formerly Sab) Sab ≡ Sab' g'b'b = (∂xa/∂x'b) // g' pulls the second index up Sab ≡ gaa'Sa'b = (∂xa/∂x'b) // g pulls the first index down Sab ≡ gaa'Sa'b' g' b'b = (∂xa/∂x'b) // both actions at once Sab ≡ gaa'Sa'b' g' b'b // the inverse of the previous line (using gabgbc = δac twice) (q) Theorem: Sab = Rba and Sab = Rba ( reflect indices in vertical line between them) This theorem has already been proven in section (i) Theorem 3 as part of the discussion there of the fact that, when tilted matrix forms are consider for R and S, one has all of the following matrix results: S-1 = ST SST = 1 STS= 1 S = RT S = R-1 R-1 = RT RRT = 1 RTR= 1 R = ST R = S-1 Here two slightly different lower-level proofs of this theorem that Sab = Rba . Once this is established, one ran raise and lower indices on either side to get all of the following Sab = Rba Sab = Rba Sab = Rba Sab = Rba As noted earlier, an implication is that one can completely eliminate references to matrix S in tensor analysis and that is what is usually done! Proof of Theorem: This proof is a bit long-winded, but brings in many earlier results: δba" Saa" = Sab // introduce a δ . Remember all g's are symmetric. (g' bb' g' a"b') Saa" = Sab // since g'ab and g'ab are inverses of each other. g' bb' δb'b" Saa" g' a"b"= Sab // reorder and introduce another δ g' bb' (Rb'a' Sa'b") Saa" g'a"b"= Sab // 1 = RS so δb'b" = (Rb'a' Sa'b") g' bb' Rb'a' (Saa"Sa'b" g'a"b") = Sab // regroup g' bb' Rb'a' (gaa') = Sab // use gaa' = Saa"Sa'b" g'a"b", see end of (i) above (g' bb' Rb'a' gaa') = Sab // regroup Rba = Sab // g and g' raise and lower R's indices, see (n) above Notice that the above theorem says Sab = (∂xa/∂x'b) = (∂x'b/∂xa) = Rba A faster way to derive this result is to differentiate dxcdxc = dx'cdx'c and use the chain rule: dx'b = ( ∂( dx'cdx'c)/∂x'b ) = (∂(dxcdxc)/∂xa) (∂xa/∂x'b) = dxa (∂xa/∂x'b) => (∂x'b/∂xa) = (∂xa/∂x'b) => Rba = Sab Similar results can be derived for other index positions (or we can just raise and lower indices!) to get Sab = Rba = (∂xa/∂x'b) = (∂x'b/∂xa) Sab = Rba = (∂xa/∂x'b) = (∂x'b/∂xa) Sab = Rba = (∂xa/∂x'b) = (∂x'b/∂xa) Sab = Rba = (∂xa/∂x'b) = (∂x'b/∂xa) Here index a is always in x-space, while index b is in x'-space. The two vector transformation rules V'a = RabVb V'a = SbaVb can now be written V'a = RabVb V'a = RabVb which has the advantage that the indices are properly arranged for matrix multiplication in both cases. Here then is a restatement of the transformation of the example given in section (l) , T'abcde = Raa' Rbb' Rcc' Sd'd Se'e Ta'b'c'd'e' // no pairs contracted becomes T'abcde = Raa' Rbb' Rcc' Rdd' Ree' Ta'b'c'd'e' // no pairs contracted It is easy to remember since the second index is always the summed index and the other index has to match (up or down) the left side of the equation. (r) Orthogonality Rules The above theorem Sab = Rba can be used to eliminate S in various forms of RS = 1: SR = 1 Sab Rbc = δac Rba Rbc = δac Rba Rbc = δac Σ 1st RTST = 1 Rba Scb = δac RbaRbc = δac Rba Rbc = δac Σ 1st RS = 1 Rab Sbc = δac Rab Rcb = δac Rcb Rab = δca Σ 2nd STRT = 1 Sba Rcb = δac Rab Rcb = δac Rcb Rab = δca Σ 2nd The four results in the right column are called orthogonality rules for R. The first pair is summed on the first index, the second on the second. In section (i) it was shown that these rules are just statements of the fact that in up or down tilted standard notation R is a real-orthogonal matrix so RRT = RTR= 1. (s) The tangent and reciprocal base vectors and expansions on same Tangent and reciprocal base vectors Here are some basic translations: (en)i → (en)i // contravariant index i (n)i → (en)i // covariant index i (En)i → (en)i // contravariant index i (n)i → (en)i // covariant index i (en)i = Sin → (en)i = Sin = Rni // contravariant index i (n)i = Rni → (en)i = Rni // covariant index i (En)i = Rnkgki → (en)i = Rnkgki = Rni // contravariant index i As noted earlier, writing a vector in bold such as en is not enough to say whether the vector is contravariant or covariant. If one form or the other is intended, one must show an index up or down, even if it is just a dummy placeholder index. As examples, S = [e1, e2, e3 .... eN ] → Sij = [(e1)k, (e2)k, (e3)k .... (eN)k] R = [1, 2, 3 .... N ]T → Rij = [(e1)k, (e2)k, (e3)k .... (eN)k]T The relationship between en and en is very simple, En ≡ g'ni ei → en = g'ni ei and en = g'ni ei For either contravariant or covariant indices (indices are not shown), g'ni raises the label on ei , and inverting one finds that g'ni lowers the label on ei. This fact makes things easy to remember. Using the fact that ei = ∂'ix , one has g'ni ei = g'ni ∂'ix = ∂'nx so the above line can be expressed as En ≡ g'ni ei → en = ∂'nx and en = ∂'nx The dot products are en em = 'nm → en em = g'nm = ∂'nx ∂'mx En em = δn,m → en em = δnm = ∂'nx ∂'mx En Em = g'nm → en em = g'nm = ∂'nx ∂'mx The "labels" on the base vectors behave in this dot product structure the same way that up and down "indices" behave. This is the motivation for En → en . Thus, the three final equations can be regarded as the same equation en em = g'nm where we can raise either or both indices/labels to get the other equations. For example, en em = g'nm = δnm . Inverse tangent and reciprocal base vectors Using the rules given above, g'↔ g R ↔ S en → u'n e'n → un En → U'n E'n → Un we can obtain the corresponding results for the inverse tangent and reciprocal base vectors: (u'n)i → (un)i // contravariant index i ('n)i → (un)i // covariant index i (U'n)i → (un)i // contravariant index i ('n)i → (un)i // covariant index i (u'n)i = Rin → (u'n)i = Rin = Sni // contravariant index i (')i = Sni → (un)i = Sni // covariant index i (U'n)i = Snkg'ki → (un)i = Snkg'ki = Sni // contravariant index i R = [u'1, u'2, u'3 .... u'N ] → Rij = [(u'1)k, (u'2)k, (u'3)k .... (u'N)k] S = ['1, '2, '3 .... 'N ]T → Sij = [('1)k, ('2)k, ('3)k .... ('N)k]T U'n ≡ gni u'n → u'n = gni u'i and u'n = gni u'i u'n u'm = nm → u'n u'm = gnm U'n u'm = δn,m → u'n u'm = δnm U'n U'm = gnm → u'n u'm = gnm Summary table. The summary table given at the end of Section 6 (e) was this x'-space x-space axis-aligned basis vectors e'n un (e'n)i= δn,i (un)i= δn,i dual partners to the above E'n Un (E'n)i = g'ni (Un)i = gni tangent base vectors u'n en (u'n)i= Rin (en)i = Sin reciprocal base vectors U'n En (U'n)i = g'ia Sna (En)i = gia Rna = gnaRia = g'naSia which translates into this → : x'-space x-space axis-aligned basis vectors e'n un (e'n)i= δni (un)i= δni dual partners to the above e'n un (e'n)i = g'ni (un)i = gni tangent base vectors u'n en (u'n)i = Rin (en)i=Sin= Rni reciprocal base vectors u'n en (u'n)i = g'ia Sna (en)i = gia Rna (u'n)i = Sni (en)i = Rni x-space expansions The x-space expansions of Section 6 (f) were V = V1 u1 + V2 u2 +... = ΣnVn un where Un V = Vn Un = gni ui V = 1 U1 + 2 U2 +... = Σnn Un where un V = n V = V'1 e1 + V'2 e2 +... = Σn V'n en where En V = V'n En = g'ni ei V = '1 E1 + '2 E2 +... = Σn 'n En where en V = 'n and they now become → : V = V1 u1 + V2 u2 +... = ΣnVn un where un V = Vn un = gni ui V = V1 u1 + V2 u2 +... = ΣnVn un where un V = Vn V = V'1e1 + V'2e2 +... = Σn V'n en where en V = V'n en = g'ni ei V = V'1e1 + V'2 e2 +... = Σn V'n en where en V = V'n x'-space expansions Similarly, the x'-space expansions of Section 6 (f) were V' = V'1 e'1 + V'2 e'2 +... = ΣnV'n e'n where E'n V' = V'n E'n = g'ni e'i V' = '1 E'1 + '2 E'2 +... = Σn'n E'n where e'n V' = 'm V' = V1u'1 + V2u'2 +... = Σn Vn u'n where U'n V' = Vn U'n = gni u'i V' = 1U'1 + 2U'2 +... = Σn n U'n where u'n V' = n and they now become → : V' = V'1 e'1 + V'2 e'2 +... = ΣnV'n e'n where e'n V' = V'n e'n = g'ni e'i V' = V'1 e'1 + V'2 e'2 +... = ΣnV'n e'n where e'n V' = V'm V' = V1u'1 + V2u'2 +... = Σn Vn u'n where u'n V' = Vn u'n = gni u'i V' = V1u'1 + V2u'2 +... = Σn Vn u'n where u'n V' = Vn Summary of all expansions: Using implied sum notation, we can now summarize the eight expansions above, plus the unit vector expansion onto n, on just two lines : V = Vn un = Vn un = V'n en = V'n en = V'n n // x-space expansions, V'n = hnV'n V' = V'n e'n = V'n e'n = Vn u'n = Vn u'n // x'-space expansions In all cases one sees a tilted index summation where one index is a vector index and the other is a basis vector label. Half the forms shown above can be obtained from the others by just "reversing the tilt". The power of the Standard Notation makes itself felt in relations like these. Due to this tilt situation, sometimes a basis like en appearing in V = V'n en is called a "covariant basis" while the basis en appearing in V = V'n en is called a "contravariant basis". Corresponding expansions of higher rank tensors are presented in section (w) below. If V is a tensor density of weight W (see Appendix D and E) the rule for adjusting the above expansions is to make the replacement V'n → JW V'n and V'n → JW V'n where J is the Jacobian of Section 5 (k). (t) Comment on Covariant versus Contravariant Consider this expansion for a vector V in x-space, V = Vncn Vn = V cn where cn is some basis having dual basis cn where as usual cn cm = δnm. Imagine taking Vn → V'n = Rnm Vm and ci → ci' = Qij cj. What Q would cause the following to be true? V = Vncn = V'nc'n In other words, how does one transform that basis cn such that the vector V remains unchanged if Vn is transformed contravariantly? The answer to this question is that Qij = Rij since then (using R orthogonality as in section (r)) V'nc'n = [Rnm Vm][ Rnj cj] = (Rnm Rnj) Vm cj = δmj Vm cj = Vj cj = Vncn Compare then the transformation of Vn with that of the basis cn: V'n = Rnm Vm cn' = Rnm cm The Vm vector components transform with Rnm but the basis vectors have to transform with Rnm to maintain the invariance of the vector V. One varies with the down-tilt R, while the other varies with the up-tilt R, so the two objects are varying against each other in this tilt sense. They are "contra-varying", so one refers to the components Vm as contravariant components with respect to the basis cm . If one starts over with Vn components and the cn "dual" (reciprocal) expansion vectors and asks for a solution to this corresponding problem, V = Vncn = V'nc'n one finds not surprisingly that the dual basis must vary as cn' = Rnm cm and then one has V'n = Rnm Vm cn' = Rnm cm which is the previous result with all indices up↔down. Comparing the tilts, one would say that the Vm again "contra vary" with the way the cm vary to maintain invariance of V. But one does not care about the dual basis, one cares about the basis, so relative to the basis cn one has V'n = Rnm Vm cn' = Rnm cm If the basis cm is varied as shown here, then the dual basis cm varies as shown above and V remains invariant. Comparing now the way the Vn transform with the way the basis vectors cm transform, one sees that both equations have the same tilted Rnm. They are "co-varying", so one refers to the components Vm as covariant components with respect to the basis cm . (u) The Significance of Tensor Analysis "Why is tensor analysis important?", the reader might ask in the midst of this storm of index shuffling. Now is a good time to answer the question. Consider the following sample equation in x-space, where the fields Q, H, T and B may or may not be tensor fields: Qadc(x) = Hab(x)Tbc(x) Bd(x) Notice that when contracted indices are ignored, the remaining indices have the same type on both sides. If the various objects really were tensors, one would say this was a "valid tensor equation" based on the index structure just described. One says that an equation is "covariant with respect to transformation x' = F(x)" if the equation has exactly the same form in x'-space that it has in x-space , which for our example would be Q'adc(x') = H'ab(x')T'bc(x') B'd(x') Here the word "covariant" has a new meaning, different from its being a type of vector or index. The meaning is related in the sense that, comparing the above two equations, everything has "moved" in the same manner ("co-varied") under the transformation. (Some authors think the word "invariant" is more appropriate; Ricci and Levi-Civita used the term "absolute".) If the objects Q, H, T and B are tensors under F, then covariance of any valid tensor equation like the one shown above is guaranteed!! The reason is that, once the contracted indices on the two sides are ignored according to the "contraction neutralization rule", the objects on the two sides of the equation have the same indices which are of the same type, so both sides are tensors of the same type, and therefore both sides transform from x-space to x'-space in the same way. If one starts, for example, with the primed equation and installs the known transformations for all the pieces, one ends up with the unprimed equation. If this explanation is not convincing, a brute force demonstration can perhaps help out. The following is also a good exercise is using the two tilt forms of the R matrix. Recall from section (q) that Sba = Rab and that SR = 1 is replaced by the various orthogonality rules of section (r). We shall process the primed equation into the unprimed one, being careful to give new summation indices unique names: Q'adc(x') = H'ab(x')T'bc(x') B'd(x') // x'-space equation [Raa'Rdd'Rcc'Qa'd'c'(x)] = [Raa'Rbb' Ha'b'(x)] [Rbb"Rcc' Tb"c'(x) ] [Rdd'Bd'(x)] = Raa' Rdd' Rcc'(Rbb' Rbb") Ha'b'(x) Tb"c'(x)Bd'(x) where we have reordered the R factors to (1) match the LHS as much as possible, and (2) to group any pairs that came from a contracted index. Using one of the orthogonality rules of section (r), = Raa' Rdd' Rcc'(δb'b") Ha'b'(x) Tb"c'(x) Bd'(x) = Raa' Rdd' Rcc' Ha'b'(x) Tb'c'(x) Bd'(x) so that (Raa'Rdd'Rcc') Qa'd'c'(x) = (Raa' Rdd' Rcc') Ha'b'(x) Tb'c'(x) Bd'(x) Now apply to both sides the factor RaA RdD RcC and sum on a,c,d. RaA RdD RcC (Raa'Rdd'Rcc') Qadc(x)] = RaA RdD RcC (Raa'Rdd'Rcc') Ha'b'(x) Tb'c'(x) Bd'(x) or (RaA Raa')( RdD Rdd')( RcC Rcc') Qadc(x) = (RaA Raa')( RdD Rdd')( RcC Rcc') Ha'b'(x) Tb'c'(x) Bd'(x) Then use a section (r) orthogonality rule in each (R R) to get (δAa')( δDd')(δCc') Qa'd'c'(x) = (δAa')( δDd')(δCc') Ha'b'(x) Tb'c'(x) Bd'(x) or QADC(x) = HAb'(x) Tb'C(x) BD(x) Finally, rename the indices A,D,C,b' to be a,d,c,b Qadc(x) = Hab(x) Tbc(x) Bd(x) // x-space equation Thus it has been shown that, if all the objects transform as tensors, the equation is covariant. Tensor density equations are also covariant. As discussed in Appendix D, a tensor density of weight W is a generalization of a tensor which has the same transformation rule as a regular tensor, but there is an extra factor of J-W on the right hand side of the rule, where J is the Jacobian J = detS. For example, Q'adc(x') = J-WQ Raa'Rdd'Rcc'Qa'd'c'(x) would indicate that Q' was a tensor density of weight WQ. If WQ = 0, then Q is a regular tensor. With this definition in mind, it is easy to generalize the notion of a "covariant equation" to include tensor densities. Consider some arbitrary tensor equation which we represent by our example above, Qadc(x) = Hab(x)Tbc(x) Bd(x) Suppose all four objects Q, H, T, B are tensor densities with weights WQ, WH, WT, WB. If the four objects Q, H, T, B are tensor densities, and if the up/down free indices match on both sides (the non-contracted indices), and if WQ = WH + WT + WB, then this is a "valid tensor density equation" and covariance is guaranteed, so it follows that Q'adc(x') = H'ab(x')T'bc(x') B'd(x') . It is trivial to edit the above proof by just adding weight factors in the right places and then of course they cancel out on the two sides. Examples of covariant tensor equations: In special relativity, which happens to involve linear Lorentz transformations, a fundamental principle is that any "equation of motion" describing anything at all (particles, EM fields, etc) must be covariant with respect to Lorentz transformations, or it cannot be a valid equation of motion (ignoring general relativity). An equation of motion must look the same in a reference frame which is rotated, boosted, or related by any combination of boosts and rotations to some original frame of reference (see Section 5 (m)). As was noted earlier, the tradition is to write 4-vector indices as Greek letters and 3-vector spatial indices as Latin letters. For example, we can define the "electromagnetic field-strength tensor" (rank-2) this way in terms of the 4-vector "vector potential" Aμ: Fμν ≡ ∂μAν - ∂νAμ where ∂μ means gμα∂α, the contravariant form of the gradient operator. The components are then where c is the speed of light and of course E and B are the electric and magnetic fields. Two of Maxwell's equations are (in SI units where ε0μ0= 1/c2) ∂νFμν = μ0 Jμ Jμ = (cρ,J) while the other two are ∂αFμν + ∂μFνα + ∂νFαμ = 0 or ∂αFμν + cyclic = 0 One can see that each of these equations involves only tensors and we expect that in x'-space these equations will take the form ∂'νF'μν = μ0 J'μ J'μ = (cρ',J') ∂'αF'μν + ∂'μF'να + ∂'νF'αμ = 0 or ∂'αF'μν + cyclic = 0 Objects like ∂νFμν and ∂αFμν are true rank-3 tensors because the transformation F is linear. The next section shows the pain encountered when F is not linear. (v) The Christoffel Business: covariant derivatives When a transformation F is non-linear, the matrix Rab is a function of x. Thus one gets the following transformation for a lower index derivative of a covariant vector field component ∂aVb(x), where a "second term" quite logically appears, (∂'aV'b) = (Rad∂d) (RbcVc) = Rad Rbc (∂dVc) + Rad(∂d Rbc)Vc This second term did not arise earlier when we looked at ∂a on a scalar field φ'(x') = φ(x) , (∂'aφ') = (Rad∂d) φ = Rad (∂d φ) In special relativity, for example, where transformations are linear, ∂d Rbc = 0, there is no second term, and the object ∂aVb transforms as a covariant rank-2 tensor, (∂'aV'b) = Rad Rbc (∂dVc) , // F is a linear transformation but in the general case the second term is present, so ∂dVc fails to transform as a rank-2 covariant tensor. In this case, one defines a certain "covariant derivative" which itself has an extra piece Vb;a ≡ ∂aVb – Γkab Vk => Vd;c ≡ ∂cVd – Γkcd Vk where Γcab is a certain function of the metric tensor g. One then finds that V'b;a = Rad Rbc V d;c or [∂'aV'b – Γ 'kab V'k] = Rad Rbc [∂dVc – Γkcd Vk] (*) so that this covariant derivative of a covariant vector field Vc transforms as a covariant rank-2 tensor even with non-linear transformation F (see Christoffel Ref., 1869). This issue arises in general relativity and elsewhere. The object Γcab (the "Christoffel connection") is given by Γcab ≡ {ab,c} ≡ ≡ gcd [ab,d] = ½ gcd( ∂agbd + ∂bgad – ∂dgab ) // Christoffel 2nd kind Γdab ≡ [ab,d] ≡ ½ ( ∂agbd + ∂bgad – ∂dgab ) // Christoffel 1st kind and this is where the various "Christoffel symbols" come into play. In general relativity, Γcab is also known as the "affine connection" which represents the effect of "curved space" appearing as a force which acts on a mass (that is to say, a gravitational force), see Section 5(n). The derivative of any tensor field other than a scalar field shows this same complication when the underlying transformation F is non-linear. For example, ∂agbd(x) does not transform as a rank-3 tensor, ∂'ag'bd(x') = (Rad∂d)(Rbb'Rdd'gb'd') = Rbb'Rdd'(∂d gb'd') + other terms and therefore neither of the Christoffel symbols Γdab or Γcab transforms as a tensor in this case. Verifying the claim (*) takes a bit of algebra which is left to the reader (see Weinberg p 103). There are various other notations for the covariant derivative, such as aBb and Bb,a . (w) Expansions of higher order tensors Appendix E clarifies the use of direct product and polyadic notations for describing the basis vector combinations onto which higher order tensors can be expanded in a simple generalization of the vector expansions presented in section (s) above. There it was shown that a vector A can be expanded in two interesting ways : A = Σn An un An are the contravariant components of A in x-space A = Σn A'n en A'n are the contravariant components of A in x'-space In the first, un are axis aligned basis vectors, and in the second en are the tangent base vectors. If A were instead a tensor of rank n, these expansions would be replaced by A = Σijk... Aijk... (uiujuk...) Aijk... are the contravariant components of A in x-space A = Σijk... A'ijk... (eiejek...) A'ijk... are the contravariant components of A in x'-space where there are n indices in each sum, n factors in the direct products, and n contravariant indices on the components of tensors A in x-space and in x'-space. In the polyadic notation the direct-product crosses are eliminated giving A = Σijk... Aijk... uiujuk... Aijk... are the contravariant components of A in x-space A = Σijk... A'ijk... eiejek... A'ijk... are the contravariant components of A in x'-space . In the case of rank-2 tensors, a product like uiuj = uiuj is called a dyadic (see App. E). In this case (only) the product can be visualized as uiuTj which is a matrix constructed from a column vector to the left of a row vector. Thus one can write A = Σij Aij uiuTj Aij are the contravariant components of A in x-space A = Σij A'ij eieTj. A'ij are the contravariant components of A in x'-space . Appendix E encourages the interpretation of a rank-2 tensor A as an operator in a Hilbert space, where the matrices Aij and A'ij are matrices associated with the operator A in different bases, Anm = <un | A | um > = the x-space components of tensor A A'nm = <en | A | em > = the x'-space components of tensor A As Appendix E shows, these two matrices are related to each other by a similarity transformation A' = R A R-1. Appendix F applies the notation shown above to the problem of writing the dyadic product (v) in arbitrary curvilinear coordinates. One can think of (v) as a matrix which acts on some vector to its right. In Cartesian coordinates one has (v)ij ≡ (∂jvi). Since (v) is not a differential operator acting on whatever vector lies to its right, we did not include an analysis of (v) in Section 10 on the gradient. The main result of Appendix F is this: (x-space is Cartesian, x'-space is that of the curvilinear coordinates): (v) = Σij [(v)(u)]ij uiujT [(v)(u)]ij = (∂jvi) (v) = Σij [(v)(e)]ij eiejT [(v)(e)]ij = Σab g'ja [g'ib (∂'av'b) + Σd Rid (∂'aRbd) v'b ] Notice that the expression for the second matrix [(v)(e)]ij contains only x'-space references ( R(x') , vb'(x'), g'(x'), ∂'). Again, one can think of (v) as an operator in a Hilbert Space and the two matrices shown above are then associated with this operator in the ui and ei bases, [(v)(u)]ij = <un |(v) | um > [(v)(e)]ij = <en |(v) | em > The object (v) looks like a rank-2 tensor but of course it is not a rank-2 tensor for the reason discussed in section (v). The complicated second term shown above in [(v)(e)]ij is similar to that "second term" which appears in the expansion of (∂'aV'b) at the start of section (v). Since (v) is not a tensor, the coefficients shown in the above expansions cannot be interpreted as the x-space and x'-space components of a tensor. Nevertheless, being able to express (v) in terms of curvilinear coordinates finds great usefulness in continuum mechanics. 8. Transformation of Differential Length, Area and Volume This Section and all remaining Sections use the Standard Notation introduced in Section 7. The term N-piped is short for N dimensional parallelepiped. The context is Picture B: Since this Section is quite lengthy, a brief overview is in order: Overview The transformation of differential length, area, and volume is first framed in terms of the mapping of an orthogonal differential N-piped in x'-space to a skewed differential N-piped in x-space. The N-piped in x'-space has axis-aligned edges of length dx'n, while the N-piped in x-space has edges endx'n where en are the tangent base vectors introduced in Section 3. We want to learn what happens to the edges, face areas and volume as one differential N-piped is mapped into the other by the curvilinear coordinate transformation x' = F(x). After solving this problem, we go on to consider the transformations of arbitrary differential vectors, areas and volume. Section (a): The differential N-piped mapping The differential N-piped mapping is described and various symbols are defined. Section (b): Properties of the finite N-piped spanned by the en in x-space Results from Appendix B concerning finite N-piped geometric properties are quoted. Certain definition changes are made to make the formulas suitable for tensor analysis. The purpose of the lengthy Appendix B is to lend credence to the general formulas for elements of area and volume in N dimensions. Section (c): Back to the differential N-piped mapping: how edges, areas and volume transform Setup. The finite N-piped edges en are scaled by curvilinear coordinate variations dx'n to create a differential N-piped in x-space having edges (endx'n). Edge Transformation. The edges en of the x-space N-piped map into axis-aligned edges e'n in x'-space. Area Transformation. Tensor density notions as presented in Appendix D are used here. Volume Transformation. The volume transformation is computed several different ways. Covariant Magnitudes. These are | dx'(n)|, | dA'n | and | dV' | in the Curvilinear View of x'-space. Two Theorems. (1) g' / h'n2 = g'nn g' = cof(g'nn) and (2) |(Πxi≠nei)| = . Cartesian-View Magnitude Ratios. Appropriate for the continuum mechanics application. Nested Cofactor Formulas: Transformation of N-piped areas of dimension N-2 and lower. Transformation of arbitrary differential vectors, areas and volume. Having built confidence with the general form of vector area and volume expressions in the N-piped case, the N-piped is jettisoned and formulas for the transformation of arbitrary vectors, areas and volume are derived. Concatenation of Transformations. What happens to the transformation of vectors, areas and volumes when two transformations are concatenated? One result is that J = J1J2. Examples of area magnitude transformation for N = 2,3,4 Example 2: Spherical Coordinates: area patches Section (d): Transformation of Differential Volume applied to Integration The volume transformation obtained in Section (c) is related to the traditional notion of the Jacobian changing the "measure" in an integration. The "Jacobian Integration Rule" can then be expressed as a distributional equation. Section (e): Interpretations of the Jacobian (a) The differential N-piped mapping Section 3 (a) above considered the following situation (dx'n > 0) : dx'(n) = e'n dx'n x'-space axis-aligned differential vector, and (e'n)i = δni dx(n) = en dx'n x-space mapping of the above vector under F-1 or R-1 dx'(n) = R(x) dx(n) relation of the two differential vectors (contravariant rule) A superscript (n) on the differentials makes clear there is no implied sum on n. The vectors dx(n) span a differential N-piped in x-space, while the dx'(n) span a corresponding differential N-piped in x'-space. The two N-pipeds are related by the mapping x' = F(x). Since the regions are differentially small, this mapping is the same as the linearized mapping dx' = R dx. The metric tensor in x-space will be taken to be g = 1, so it is a Cartesian space. As discussed at the end of Section 5 (a), the x'-space N-piped can be viewed in (at least) two ways depending on how the metric tensor g' is set. For a continuum mechanics flow application, one sets g' = 1 and this gives the Cartesian View of the x'-space N-piped. For such flows dot products and magnitudes of vectors like dx(n) are not invariant under the transformation. For our curvilinear coordinates application, however, we set g' = RgRT = RRT and this causes vector dot products and magnitudes to be invariant and we can talk about such objects as being tensorial scalars. This is the Curvilinear View of x'-space. When g' ≠ 1, it is impossible to accurately represent the Curvilinear-View picture of the x'-space N-piped as a drawing in physical space (for N=3). This subject is discussed for a sample 2D system in Appendix C (e). Although the basis vectors (e'n)i = δni in x'-space are always axis-aligned, they are only orthogonal for an orthogonal coordinate system, since e'ne'm = enem = g'nm . Nevertheless, even for a non-orthogonal system we draw the axes as if they were orthogonal, which at least provides a representation of the notion of "axis aligned" basis vectors. For N > 3 one at least imagine this kind of drawing. Due to these graphical difficulties, in the drawing below the Cartesian View of x'-space is shown. Since g' = 1 for this situation, the axis-aligned basis vectors e'n are in fact unit vectors 'n and are orthogonal, so the picture becomes at least comprehensible: The orthogonal Cartesian-View N-piped allows visualization of these curvilinear coordinate variations, all dx'k > 0, dL'n ≡ dx'n dA'n ≡ Πi≠ndx'i dV' ≡ Πidx'i = dAn dLn For example, for N=3 one would have dL'1 ≡ dx'1 dA'3 = dx'1dx'2 dA'1 = dx'2dx'3 dA'2 = dx'3dx'1 dV' = dx'1dx'2dx'3 The Cartesian-View x'-space N-piped is always orthogonal because the (e'n) are orthonormal axis-aligned unit vectors (since g'=1). In contrast, the x-space N-piped is typically rotated and possibly skewed as well (if the coordinates x'i describe a non-orthogonal coordinate system). The transformation F and its linearized version R map the skewed x-space N-piped into the orthogonal x'-space N-piped. As one moves around in x-space so that point x changes, the picture on the left above keeps its shape, just translating itself to the new point x', but the picture on the right changes shape and volume because the vectors en(x) are functions of x. It is our goal to write expressions for edges, areas and volumes in these two spaces and to then show how these objects transform between the two spaces. To this end, we shall rely on work done in Appendix B which is summarized in the next section. Following that, we shall add to each edge a differential associated with that edge (such as er → erdr in spherical coordinates), and that will bring us back to the differential N-piped picture above. (b) Properties of the finite N-piped spanned by the en in x-space The finite N-piped spanned by the tangent base vectors en in x-space has the following properties (as shown in Appendix B) : The N spanning edges are the vectors en which have lengths |en| = h'n (scale factors ). There are 2N vertices. There are N pairs of faces. The two faces of each pair are parallel in N dimensions. One face of each pair touches the point where the tails of all the en vectors meet (the near face) while the other does not touch this meeting point (the far face). Each face of an N-piped is an (N-1)-piped having 2N-1 vertices. The faces are planar surfaces of dimension N-1 embedded in an N dimensional space. A face's vector area An is spanned by all the ei except en and is labeled by this missing en vector. The far face has out-facing vector area An , and the near face has out-facing area vector – An. These vector areas are normal to the faces. The vector area An is given by several equivalent expressions: An = |det(Sab)| en An = σ (-1)n-1 Πxi≠n ei An = σ (-1)n-1 e1 x e2 ... x eN // en missing σ ≡ sign[det(Sab)] = sign[det(Rab)] (An)i = σ (-1)n-1 εiabc..x (e1)a(e2)b.... (eN)x // en missing The volume of the N-piped is given by (see Section 5 (k) concerning J) V = | det [ e1, e2, e3 ... eN] | = | det(Sab) | = g'1/2 = |J| The mapping picture above does not apply to a finite N-piped. The finite N-piped just discussed exists in x-space. One might ponder into what shape it maps in x'-space under the transformation F. In the case of spherical coordinates (Section 1 Example 1), all of x-space maps into a certain orthogonal "office building" in x'-space. A finite N-piped within x-space maps into some very complicated 3D shape within the office building which is bounded by curves which in general are not even coplanar. The point is that a finite N-piped in x-space does NOT map back into some nice orthogonal N-piped in x'-space and the picture drawn in the previous section does not apply. However, when differentials are added in the next section, then, since the mapped regions are very small, the mapping of the x-space differential N-piped is in fact an "orthogonal" N-piped in x'-space. This is because for a tiny region near some point x, the mapping between dx and dx' is described by matrix R. Conventions for defining the area vector and volume. In Appendix B (and as shown above) the area vector An is defined so that the out-facing normal of the x-space N-piped's "far face n" is An, regardless of the sign of det(S). This was done to simplify the computation of the flux of a vector field emerging from the N-piped in the geometric divergence calculation in Section 12. That calculation is then valid for either sign of det(S), a sign that we call σ. The following picture illustrates on the left an x-space 3-piped which is obtained by reverse-mapping the x'-space orthogonal 3-piped using an S which has det(S) > 0. On the right one sees the x-space 3-piped that results for S → -S. These two 3-pipeds are related by a parity inversion of all points through the origin. If the origin lies far away, these two N-pipeds lie far away from each other, a fact not illustrated in the picture: Notice that A3 for the "far face 3" is outfacing in both cases. This definition of the vector area is not suitable for the vector analysis we are about to undertake. Instead of the above situation, we will redefine A3 = e1xe2 for both pictures, and this will cause the A3 vector in the right picture to point up into the interior of the N-piped. This new definition allows us to interpret A3 = e1xe2 as a "valid vector equation" to which we may apply the ideas of covariance and tensor densities. Notice that under a parity transformation, this newly defined A3 is a "pseudovector" which is one which does not reverse direction under a parity transformation, since A3 = (-e1) x (-e2). The subject of parity and handedness and the sign of det(S) is discussed more in Section 6 (i). Here then are the expressions for An with this new definition, where the new forms are obtained from the previous ones by multiplying by σ = sign ( det(S) ) : An = det(Sab) en = J en An = (-1)n-1 e1 x e2 ... x eN // en missing (An)i = (-1)n-1 εiabc..x (e1)a(e2)b.... (eN)x // en missing The second line shows that pseudovectors can only exist for an odd number of dimensions (such as N=3). A similar redefinition of the volume will now be done. In Appendix B the volume is defined so as to be a positive number regardless of σ with the result V = |det(S)|. We now redefine the volume by multiplication by σ, so that now V = det(S) which is of course a negative number when det(S) < 0, which in turn means the en are forming a left-handed coordinate system as per Section 6 (i). So: V = det [ e1, e2, e3 ... eN] = det(Sab) = J (c) Back to the differential N-piped mapping: how edges, areas and volume transform The Setup. If the edges of the finite N-piped described above are scaled by positive differentials dx'n > 0, the result is a differential N-piped in x-space which has the properties listed above with the following edges and areas and volume: dx(n) = en dx'n // edges dAn = J en (Πi≠ndx'i) // areas dAn = (-1)n-1 (dx'1e1) x (dx'2e2) ... x (dx'NeN) // en missing from cross product = (-1)n-1 e1 x e2 ... x eN (Πi≠ndx'i) // en missing from cross product = (-1)n-1 (Πxi≠nei) (Πi≠ndx'i) // shorthand of Appendix A (i) (dAn)i = (-1)n-1 εiabc..x (e1)a(e2)b.... (eN)x (Πi≠ndx'i) // en factor and index missing dV = det [dx'e1, dx'e2, dx'e3 ... dx'eN] = det [ e1, e2, e3 ... eN] (Πidx'i) = J det(Sab) = εabc..x (dx'e1)a(dx'e2)b..... (dx'eN)x where dAn and dV obtained from An and V according to the new definitions described above. These equations apply to the right side of the N-piped mapping picture which is replicated here For spherical coordinates, the N-piped on the right would be spanned by these vectors (Sec. 3, Ex. 2 ) erdr = dr, eθdθ= rdθ eφdφ = rsinθ dφ Edge Transformation. The edge dx(n) we know transforms as a tensorial vector under transformation F, so dx'(n) = R dx(n) where dx(n) = en dx'n Evaluation gives [dx'(n)]i = Rij (en)j dx'n = Rij Sjn dx'n = (RS)in dx'n = δin dx'n => dx'(n) = e'n dx'n since (e'n)i = δni so this contravariant edge points in the n axis direction in x'-space. Area Transformation. How does dAn transform under F? Looking at the component form stated above (dAn)i = (-1)n-1 εiabc..x (e1)a(e2)b.... (eN)x (Πi≠ndx'i) // en factor and index missing it is seen that dAn is a combination of tensor objects like εiabc..x and (e2)b. As discussed in Appendix D, a vector density of weight 0 is an ordinary vector such as e2 or dx. The ε tensor (rank-N) is a tensor density of weight -1. In forming more complicated tensor objects, the rule is that one adds the weights of the objects being combined. Therefore, one may conclude that the object dAn is a vector density of weight -1. This has the immediate implication that dAn transforms under F according to the rule dA'n = J R dAn or (dA'n)i = J Rij (dAn)j // J-W = J-(-1) = J where J = det(S) is the Jacobian of Section 5 (k). Moreover, the above equation for (dAn)i is a "valid tensor density equation" as per Section 7 (u) and is therefore covariant. This means that in x'-space the equation has the exact same form, but tensor objects are primed (the dx'i are constants), (dA'n)i = (-1)n-1 ε'iabc..x (e'1)a(e'2)b.... (e'N)x (Πi≠ndx'i) // e'n factor and index missing Insertion of ε'iabc..x = J2 εiabc..x ( Appendix D (e) 3) and (e'n)i = δni then gives (dA'n)i = (-1)n-1 J2 εiabc..x δ1aδ2b...δNx (Πi≠ndx'i) = (-1)n-1 J2 εi123..N (Πi≠ndx'i) // index n missing on ε = δni J2(Πi≠ndx'i) The last step follows from the fact that εi123..N with n missing must vanish if i ≠n, and if i=n then εn123..N = (-1)n εi123..N = (-1)n. The conclusion then is that dA'n = J2(Πi≠ndx'i) e'n since (e'n)i = δni // Section 7 (s) and the covariant vector area dA'n points in the n-axis direction in x'-space. For the reader dubious of the claim that dA'n = J R dAn, consider: (dA'n)i = J Rij (dAn)j = J Rij {[ (-1)n-1 εiabc..x (e1)a(e2)b.... (eN)x (Πi≠ndx'i) } // en missing = J Rij {[ (-1)n-1 εiabc..x R1a R2b.... RNx (Πi≠ndx'i) } // Rnκ missing = J {[ (-1)n-1 εiabc..x Rij R1a R2b.... RNx (Πi≠ndx'i) } // Rnκ missing = J {[ (-1)n-1 εiabc..x Sji Sa1 Sb2.... SxN (Πi≠ndx'i) } // Sκn missing If i ≠n, one has a determinant with two columns the same since Sκn is missing so the result is 0 so the result is proportional to δni. Continuing, = δni J { (-1)n-1 εnabc..x Sjn Sa1 Sb2.... SxN (Πi≠ndx'i) } // Sκn missing in group = δni J { (-1)n-1 εnabc..x Sa1 Sb2.... Sjn.... SxN (Πi≠ndx'i) } = δni J { εabc..n...x Sa1 Sb2.... Sjn.... SxN (Πi≠ndx'i) } = δni J { det(S) (Πi≠ndx'i) } = δni J2 (Πi≠ndx'i) which agrees with the result just obtained from the covariant x'-space equation. Volume Transformation. What about the volume dV? Assume for the moment that dV is correctly represented this way, where any n will do: dV = dAn dx(n) = (dAn)i [dx(n)]i // no implied sum Installing our expression for dAn and dx(n) gives dV = J en (Πi≠ndx'i) (en dx'n) = J (Πidx'i) en en = J (Πidx'i) which is seen to agree with the modified dV definition stated above. Looking at dV = (dAn)i [dx(n)]i, dV is seen to be a tensor combination of a vector density of weight -1 with a vector density of weight 0 (an ordinary vector dx(n)), so according to Appendix D, dV must be scalar density of weight -1. This then tells us that dV' = J dV Since dV = J (Πidx'i) one gets dV' = J2 (Πidx'i) . Again, one can verify this last result from the x'-space covariant form of the equation: dV' = dA'n dx'(n) = { J2(Πi≠ndx'i) e'n } {e'n dx'n) = J2(Πidx'i) e'n e'n = J2(Πidx'i) since e'n e'n = enen = 1. An alternate derivation of the dV transform rule dV' = J dV comes from just staring at dV = εabc..x (dx'e1)a(dx'e2)b..... (dx'eN)x which by the argument above is seen directly to transform as a tensor density of weight -1. To complete the circle, we can verify for a second time the claim made above that dV = dAn dx(n) : dAn dx(n) = (dAn)i [dx(n)]i = {(-1)n-1 εiabc..x (e1)a(e2)b.... (eN)x (Πk≠ndx'k) } (dx'nen)i // en missing ..... = εabc..i..x (e1)a(e2)b.. (en)κ.. (eN)x (Πkdx'i) = det(S) (Πkdx'i) = dV To summarize the above, by examining the vector density nature of our various objects, we have been able to determine exactly how edges, areas and the volume transform under transformation F: edges dx'(n) = R dx(n) or [dx'(n)]i = Rij [dx(n)] j // ordinary vector areas dA'n = J R dAn or (dA'n)i = J Rij (dAn)j // vector density W = -1 volume V' = J dV // scalar density W = -1 Covariant Magnitudes. The x'-space magnitudes here are the "covariant" ones which are associated with the Curvilinear View of x'-space, as discussed above. Since dx(n)is a vector, it follows that | dx'(n)|2 = dx'(n) dx'(n) = dx(n) dx(n) = | dx(n)|2 => | dx'(n)| = | dx(n)| Since dAn is a vector density of weight -1, if follows that | dA'n |2 = dA'n dA'n = J2 dAn dAn = J2| dAn |2 => | dA'n | = |J| | dAn | where we note that dA'n dA'n, being a combination of two weight -1 vector densities, is a scalar density of weight -2 and thus transforms as shown above. For completeness, we can add from the above table, V' = J dV => | dV' | = |J| | dV | . Moreover it has been shown above that dx(n) = en dx'n => | dx(n)| = |en| dx'n = h'n dx'n dx'(n) = e'n dx'n => | dx'(n)| = |e'n| dx'n = h'n dx'n dAn = J (Πi≠ndx'i) en => | dAn| = |J| |en| (Πi≠ndx'i) = (|J| /h'n) (Πi≠ndx'i) dA'n = J2(Πi≠ndx'i) e'n => | dA'n| = J2 |e'n| (Πi≠ndx'i) = (|J|2/h'n) (Πi≠ndx'i) dV = J (Πidx'i) => |dV| = |J| (Πidx'i) dV' = J2 (Πidx'i) => |dV'| = |J2 (Πidx'i) which can be written more compactly using the Cartesian-View coordinate variation groupings, dx(n) = en dL'n => | dx(n)| = h'n dL'n dL'n ≡ dx'n dx'(n) = e'n dL'n => | dx'(n)| = h'n dL'n ` dAn = J dA'n en => | dAn| = (|J| /h'n) dA'n dA'n ≡ Πi≠ndx'i dA'n = J2 dA'n e'n => | dA'n| = (|J|2/h'n) dA'n dV = J dV' => |dV| = |J| dV' dV' ≡ Πidx'i dV' = J2 dV' => |dV'| = |J|2 dV' The covariant edge magnitude is of course unchanged by the transformation since it is a scalar, while the area and volume magnitudes are magnified by |J| in going from x-space to x'-space. Since dV = dAn dx(n), it is clear that the area's transformation factor of |J| is passed onto the volume. Notice that dV' is always positive regardless of the sign of J, and this is because x'-space is always a right-handed coordinate system, as in Section 6 (i). In contrast, dV can have either sign depending on the sign of J = det(S), and so dV < 0 when the en form a left-handed coordinate system in x-space. Two Theorems. We now pause to prove two small theorems which will be used below. Theorem 1: g' / h'n2 = g'nn g' = cof(g'nn) where g' = det(g'ij). The left equality is obvious since g'nn = 1/h'n2 . The right equality can be shown as follows: (g'up)ab ≡ g'ab (g'dn)ab ≡ g'ab g'up = (g'dn)-1 = cof(g'dnT)/det(g'dn) = cof(g'dn)/det(g'dn) => (g'up)nn = cof[(g'dn)nn]/det(g'dn) or g'nn = cof[g'nn] / g' QED Theorem 2: |(Πxi≠nei)| = The quantity on the left is this |(Πxi≠nei)| ≡ | e1 x e2 ... x eN | where en is missing from cross product. We shall give three quick proofs of Theorem 2, the last being valid only for N=3. First, at the start of section (c) above, one form for dAn is given by dAn = (-1)n-1 (Πxi≠nei) (Πi≠ndx'i) = (-1)n-1 (Πxi≠nei) dA'n => | dAn| = |(Πxi≠nei)| dA'n Comparison of this last result with the third line of the table above shows that the following must be true |(Πxi≠nei)| = (|J| /h'n) = g'1/2/h'n = (g' / h'n2)1/2 = QED where the last step follows from Theorem 1. Here is a more direct proof: | Πxi≠n (ei)|2 = Πxi≠n (ei) Πxj≠n (ej) = [ Πxi≠n (ei)]k [ Πxj≠n (ej)]k = [ εkabc...x (e1)a(e2)b ...... (eN)x] [ εka'b'c'...x (e1)a'(e2)b' ...... (eN)x'] // en missing in both = εkabc...x εka'b'c'...x {(e1)a(e2)b ...... (eN)x } (e1)a'(e2)b' ...... (eN)x' // en missing = e1e1 e2e2 .... eNeN + all signed permutations of the 2nd labels // en missing = g'11g'22..... g'NN + all signed permutations of the 2nd indices // en missing But this is last object is the determinant of the g'ij matrix with g'nn crossed out, which is to say, it is the minor of g'nn. Since g'nn is a diagonal element, the minor and cofactor are the same. Thus, this last object is in fact just cof(g'nn). QED. A proof for N=3 uses normal vector algebra. Setting n = 1, for example, one needs to show that | Πxi≠1 ei |2 = | e2 x e3 |2 = cof(g'11) To this end, use the vector identity (A x B) (A x B) = A2B2 – (AB)2 to show that | e2 x e3 |2 = (e2 x e3) (e2 x e3 ) = |e2|2 |e3|2 - (e2e3)2 = g'22 g'33 - (g'23)2 = cof(g'11). and the cases n = 2 and 3 are similar. Cartesian-View Magnitude Ratios. In the Cartesian View of x'-space one can write the Cartesian x'-space magnitudes as | dx'(n)|c = dL'n | dA'(n)|c = dA'n | dV'|c = dV' Then from the three x-space equations in the above table (just above Two Theorems) one obtains the following three ratios of x-space objects divided by their corresponding Cartesian-View x'-space objects: | dx(n)|/ dL'n = h'n = [g'nn]1/2 = the scale factor for edge dx(n) | dA(n)|/ dA'n = (1/h'n) |J| = (1/h'n) g'1/2 = [ g'nn g']1/2 = [cof(g'nn)]1/2 // Theorem 1 above |dV| / dV' = |J| = g'1/2 // g' ≡ det(g'ij) = J2 In these expressions, all g' references refer to the x'-space metric tensor which results in scalars actually being scalar under the transformation F: g'nm ≡ RniRmjgij = RniRmi = SinSim where Sin = (en)i . That is to say, g'nm is determined by the basis vectors en of the N-piped in x-space. In the continuum mechanics application mentioned in Section 5 (o), the actual metric tensor for x'-space is set to g' = 1 which means scalars are no longer scalars under F. The "Cartesian View" of x'-space then coincides with the physical x'-space and the ratios given above are then the physical length, area and volume magnitude ratios. The use here of a scripted g is just a temporary contrivance to allow distinction between g'nm ≡ SinSim and g'nm = δnm for this particular "non-covariant" application of tensor analysis. From now on, we restore g' to its usual meaning g'nm ≡ RniRmjgij . It is convenient to make the definition dAn ≡ | dA(n)| and then the above area magnitude ratio relation may be written dAn = dA'n and dA'n = Πi≠ndx'i is just a product of the appropriate curvilinear coordinate variations. Nested Cofactor Formulas: The object dAn is the "area" of a face on an N-piped. This face, which is itself an (N-1)-piped, in turn has its own "areas" which are (N-2)-pipeds, and so on, so there is a hierarchy of "areas" of dimensions N-1 all the way down. The area ratios of corresponding areas under transformation F are determined by equations similar to that above. For example, the mth face of face n of an N-piped has area ratio . This at least makes some sense since the matrix cof(g'nn) has dimension N-1, so its cofactor matrix has dimension N-2, and so on. For N=3, these faces would be line segments and one would have cof(g'33) = => cof[cof(g'33)]22 = g'11 = h'12 => = h'1 and h'1 is in fact the edge length ratio given above. Transformation of arbitrary differential vectors, areas and volume. The above discussion is geared to the N-piped transform picture and reveals how dx(n), dA(n) and dV transform where all these quantities are directly associated with the particular differential N-piped in x-space spanned by the (endx'n) vectors. But suppose dx is an arbitrary differential vector in x-space. Certainly dx' = R dx, so we know how this dx transforms under F. Based on the work in Appendix B as carried through into the N-piped transform discussion above, it seems clear that an arbitrary differential area in x-space dA can be represented as follows : dA = (dx[1]) x (dx[2]) ... x (dx[N-1]) (dA)i = εiabc..x (dx[1])a(dx[2])b.... (dx[N-1])x , where the dx[i] are an arbitrary set of N-1 linearly independent differential vectors in x-space. For N=3 one would write dA = dx[1] x dx[2]. Since ε has weight -1 and all other objects are true vectors (weight 0), one again concludes that dA transforms as a tensor density of weight -1, so dA' = J R dA. If any of these dx[k] were a linear combination of the N-2 others, one could say dx[k] = Σjαkj dx[j] and then the above (dA)i expression would give a sum of terms each of which vanishes by symmetry, resulting in (dA)i = 0. Finally, given the set of dx[i] linearly independent vectors shown above for i = 1,2..N-1, we can certainly find one more such that all N are then linearly independent, so then we have a set of N arbitrary differential vectors dx[i] ( arbitrary as long as they are linearly independent), and these will form a volume in x-space, dV = det (dx[1], dx[2], ..... dx[N] ) = εabc..y (dx[1])a(dx[2])b.... (dx[N])y By the argument just given, inspection shows that this dV transforms as a tensor density of weight -1 so dV' = JdV. These then are the transformation rules for arbitrary differential vectors, areas and volumes transforming under F : dx' = R dx |dx'| = |dx| dA' = J R dA |dA'| = |J| |dA| dV' = JdV |dV'| = |J| |dV| J = det(S) J2 = g' To review, then, in Cartesian x-space we have these expressions for area and volume dA = (dx[1]) x (dx[2]) ... x (dx[N-1]) dV = det [dx[1], dx[2], ... dx[N]] . Written out in terms of covariant vector components these expressions appear as, (dA)i = εiabc..x (dx[1])a(dx[2])b.... (dx[N-1])x dV = εabc..y (dx[1])a(dx[2])b.... (dx[N])y , where ε is the usual permutation tensor involved in cross products and determinants. As noted earlier, these are both "valid tensor density equations" (end of Section 7 (u)) , so they can be written in x'-space as (dA ')i = ε'iabc..x (dx'[1])a(dx'[2])b.... (dx'[N-1])x dV' = ε'abc..y (dx'[1])a(dx'[2])b.... (dx'[N])y , where (App. D) each ε' can be written as ε' = J2ε = g' ε, and where each dx' = Rdx. Therefore: (dA ')i = g' εiabc..x (dx'[1])a(dx'[2])b.... (dx'[N-1])x dV' = g' εabc..y (dx'[1])a(dx'[2])b.... (dx'[N])y . Since the permutation tensor now appears, these can be written in terms of cross products, dA' = g' (dx'[1]) x (dx'[2]) ... x (dx'[N-1]) dV' = g' det [dx'[1], dx'[2], ... dx'[N]] . Using contravariant components [dx'[1]]i, this cross product and determinant are formed just as they are in x-space. These two equations then give at least some feel for "the meaning of dA' and dV' in x'-space" . Going back to x-space, we could have written the equations there using g = det(g'ij) = det(δij) = 1 : dA = g (dx[1]) x (dx[2]) ... x (dx[N-1]) dV = g det (dx[1], dx[2], ... dx[N]) . This then is a useful intepretation of the covariant form of these equations. Adding primes to everything in the above two equations yields the previous two equations and only the permutation tensor ε is involved in both sets of equations. With this understanding, the above transformation rules for differential vectors, areas and volumes can be extended from Picture B to the more general Picture A, where g is an arbitrary metric tensor, Concatenation of Transformations. Consider x" = F2(x') and x' = F1(x) so that x" = F2[F1(x)] ≡ F(x). At some point x, the linearization will yield dx" = R2R1dx as the rule for vector transformation, so the R matrix associated with transformation F is R = R2R1, and then S = R-1 = R1-1R2-1 = S1S2. Each transformation will have an associated Jacobian: J1 = det(S1) and J2 = det(S2). The concatenated transformation then has J = det(S) = det(S1S2) = det(S1)det(S2) = J1J2. The implication is that when one concatenates two transformations in this manner, the area and volume transformations shown above still apply, where J is taken to be the product of the two underlying Jacobians. For example, one could consider the mapping between two different skewed N-pipeds, each representing a different curvilinear coordinate system, with our orthogonal N-piped as an intermediary object, In this case one has F = F2-1F1, so R = R2-1R1 = S2R1 and then S = S1R2 so J = J1/J2. This J then would be used in the above area and volume transformation rules, for example, dAleft = J R dAright. In the continuum mechanics flow application, g = 1 on both left and right as well as center, time t0 is on the right, time t on the left, and the volume transformation is given by dVleft = J dVright where J is associated with the combined F. This J then characterizes the volume change between an initial and final flow particle where each is skewed in some arbitrary manner. Examples of area magnitude transformation for N = 2,3,4 In the previous section it was shown that dAn = dA'n. Since this is a somewhat strange result, some examples are in order. Recall that the dAn are the areas of the faces of the differential N-piped in x-space, while the dA'n are the curvilinear coordinate variations one can visualize in the Cartesian-View picture shown above. For N=2 the area magnitude transformation results are (for a general non-orthogonal x'-space system) dA1 = dA'1 = h'2 dA'1 dA'1 = dx'1 = dL'1 dA2 = dA'2 = h'1 dA'2 dA'2 = dx'2 = dL'2 These equations are simple because the area of a parallelogram "face" is the length of an edge and so these equations just coincide with the length transformation results stated above . Remember that a face is labeled by the index of the vector which does not span the face, so h2' appears in the face 1 equation. For N=3 the area magnitude transformation results are dA1 = dA'1 dA'1 = dx'2dx'3 dA2 = dA'2 dA'2 = dx'3dx'1 dA3 = dA'3 dA'3 = dx'1dx'2 For an orthogonal N=3 system the metric tensor g'ab is diagonal, and then the above simplifies to dA1 = dA'1 = h'2 h'3 dA'1 dA'1 = dx'2dx'3 dA2 = dA'2 = h'3 h'1 dA'2 dA'2 = dx'3dx'1 dA3 = dA'3 = h'1 h'2 dA'3 dA'3 = dx'1dx'2 For an N=4 orthogonal system, dA1 = dA'1 = h'2 h'3 h'4 dA'1 dA'1 = dx'2dx'3dx'4 dA2 = dA'2 = h'1 h'3 h'4 dA'2 dA'2 = dx'3dx'4dx'1 dA3 = dA'3 = h'1 h'2 h'4 dA'3 dA'3 = dx'4dx'1dx'2 dA4 = dA'4 = h'1 h'2 h'3 dA'4 dA'4 = dx'1dx'2dx'3 Example 2: Spherical Coordinates: area patches Consider again dAn = dA'n. Since spherical coordinates are orthogonal, the orthogonal N=3 example above may be used. Example 2 of Section 5 showed that [ 1,2,3 = r,θ,φ ] h'1 = h'r = 1 dA'1 = dx'2dx'3 = dθdφ h'2 = h'θ = r dA'2 = dx'3dx'1 = drdφ h'3 = h'φ = rsinθ dA'3 = dx'1dx'2 = drdθ Therefore dA1 = dA11 => dAr = dAr r = dAr with dAr = h'2 h'3 dA'1 = r2sinθ dθdφ dA2 = dA22 => dAθ = dAθ θ = dAθ with dAθ = h'3 h'1 dA'2 = rsinθ drdφ dA3 = dA33 => dAφ = dAφ φ = dAφ with dAφ = h'1 h'2 dA'3 = rdrdθ so that dAr = r2sinθ dθdφ ρdφ rdθ ρ = rsinθ dAθ = rsinθ drdφ ρdφ dr dAφ = rdrdθ rdθ dr where all three vectors are seen to have the correct dimensions L2. As an exercise in staring, the reader is invited to verify these results from the picture below using the hints shown above on the right, (d) Transformation of Differential Volume applied to Integration As discussed in Appendix C (h), the integral ∫D dV h(x) is the same regardless of the way the dV elements are chosen, as long as those elements exactly fill the integration region D. In the discussion above, |dV| (call it dVa) refers to a positive differential volume element in x-space which is typically not aligned with the axes and for a general transformation F is not in general orthogonal. Moreover, the shape of the differential volume N-piped varies over the region of integration. Nevertheless, this "rag-tag band" of differential volumes, as noted in Appendix C for the 2D case, fills the integration region perfectly. Alternatively one could consider |dV| (call it dVb) to be the usual dx1dx2.....dxN differential volume elements, and of course this set of differential volume elements also fills the integration space perfectly. Thinking of these two different differential volumes as dVa and dVb , one can see from the definition of the integral as the limit of a sum, lim Σi dVa(xi) f(xi) = lim Σi dVb(xi) f(xi) that ∫D dVa h(x) = ∫D dVb h(x) There would be little meaning to the statement dVa = dVb, since no one is claiming there is some particular skewed N-piped of volume dVa which matches some axis-aligned N-piped of volume dVb . Nevertheless, one could write dVa = dVb as a distributional symbolic equality where the meaning of that symbolic equality is precisely the equivalence of the two integrals above for any domain D and for any reasonable function h(x). [ Formally one might have to require h(x) to be a "test function" φ(x). Certainly one would require that both integrals converge. ] What has been shown in the previous section, regarding the Jacobian, is that dVa = |J(x')| dV' = |J(x')| ( Πi=1N dx'i) |J(x')| = // g = +1 Combining this with the distributional symbolic equation dVa = dVb gives dVa = dVb |J(x')| ( Πi=1N dx'i) = ( Πi=1N dxi) or |J(x')| dV' = dVb Now overriding our previous notation, we can make these new commonly used definitions dV ≡ ( Πi=1N dxi) dV' ≡ ( Πi=1N dx'i) and express the distributional result as |J(x')| dV' = dV We refer to this distributional equality in Appendix C as the "Jacobian Integration Rule". The symbolic equation is a shorthand for this equation ∫D dV h(x) = ∫D' dV' |J(x')| h(x) where on the right h(x) = h(x(x')) and region D' is the same region as D expressed in terms of the x' coordinates. Writing out the volume elements this says ∫D ( Πi=1N dxi) h(x) = ∫D' ( Πi=1N dx'i) |J(x')| h(x(x')) and finally using Section 5 (k), ∫D ( Πi=1N dxi) h(x) = ∫D' ( Πi=1N dx'i) [ ] h(x(x')) For example, when applied to polar and spherical coordinates, one gets ∫D dxdy h(x) = ∫D' drdθ [r] h(x(r,θ)) = r ∫D dxdydz h(x) = ∫D' drdθdφ [ r2sinθ ] h(x(r,θ,φ)) = r2 sinθ In the first case h(x) = h(x,y) and h(x(r,θ)) = h(rcosθ,rsinθ). In the second case h(x) = h(x,y,z) and h(x(r,θ,φ)) = h(rsinθcosφ,rsinθsinφ,rcosθ). (e) Interpretations of the Jacobian Using Section 5 (k) facts (in standard notation) and the above sections, one can produce various expressions and interpretations for the Jacobian J and its absolute value |J| : J(x') ≡ det(Sij(x')) = det(∂xi/∂x'k) = 1/det(Rij(x(x')) = 1/ det(∂x'i/∂xk) // Section 5 (k) |J(x')| = = // Section 5 (k) |J(x')| = the volume of the N-piped in x-space spanned by the en(x), where x = F-1(x') |J(x')| = dVN-piped/dV' = ratio of differential x-space N-piped volume / ( Πi=1N dx'i) |J(x')| = dV/dV' = ( Πi=1N dxi)/ ( Πi=1N dx'i) // distributional Jacobian Integration Rule As discussed in Section 6 (i), if the curvilinear coordinates are ordered so that the en form a right handed coordinate system, then det(S)>0, σ = sign(det(S)) = +1, and |J| = J. 9. The Divergence in curvilinear coordinates (a) Geometric Derivation of the Curvilinear Divergence Formula In Cartesian coordinates div B = B = ∂nBn, but expressed in curvilinear coordinates the right side has a more complicated form. We provide here a geometric derivation of the formula for the divergence of a contravariant vector field expressed in curvilinear coordinates, which means x'-space coordinates with Picture B. This derivation is an exercise in using the transformation results obtained in Section 8 above, and in understanding the meaning of the components of a vector, as discussed in Appendix C (d). The divergence of a vector field can be computed in a Cartesian x-space by taking the limit of the flux emerging from a closed volume divided by the size of the volume, in the limit that the volume shrinks down around some point x. Being a scalar field, the divergence is a property of the vector field at some point x and therefore cannot depend on the shape of the closed volume used for the calculation. If the shape of the volume is taken to be a standard-issue axis-aligned N-piped, the divergence obtained will be expressed in terms of the Cartesian coordinates and in terms of the Cartesian components of the vector field: [div B](x) = ∂nBn(x) where B = Bn. However, if the N-piped shape is the one below, evaluation of this same [div B](x) produces an expression which involves only the curvilinear coordinates and the curvilinear components of the vector field, as will now be demonstrated. We start by considering again our differential non-orthogonal N-piped sitting in x-space, which has edges endx'n, faces dAn and volume dV, as discussed in Section 8 (c) above: In order to avoid confusion with volume V or area A, we name the vector field B. As just noted, the divergence of a vector field B is the total flux flowing out through the faces of the N-piped divided by the volume of the N-piped, in the limit that all differentials go to 0. Thus one writes symbolically, [div B](x) = (1/dV) ∫ dAB = (1/dV) ∫ dA(x)B(x) where the surface integral is over all the faces of the above x-space differential N-piped. Recall that as we move around in space, the shape (and size) of the above N-piped changes, so the dA of a face changes, hence dA(x). Comment on the div B as a scalar. If B is a tensorial vector field, then div B is a tensorial scalar field, and one can write [div B]'(x') = [div B](x) The operator object (1/dV)∫dA(x) acts as a tensorial vector operator so that the result of its action on B is a tensorial scalar. In Section 8 it was shown that dA is a vector density of weight -1 and so is dV. This means that dV' = J dV and dA' = J RdA so the ratio dA/dV is a tensorial vector. The fact that div B is a tensorial scalar is more obvious from the alternative divergence derivation given in section (e) below. The task is now to compute the integral ∫dA(x)B(x). Appendix B shows that the N-piped faces come in parallel pairs, so we start by considering pair n. As shown in Section 8 (c) , the vector area of the far face of pair n is given by dAn(x) = |det(Sij(x'))| en(x) ( Πi≠n dx'i) = en(x) ( Πi≠n dx'i) Here |det(S)| = |J| = g'1/2 (Section 5 (k)), and en are the reciprocal base vectors (Section 6) . Quantities dAn, S, en and g' are explicitly shown as functions of space. Appendix B (c) shows that the out-facing vector area for the far face of pair n is dAn, while the out-facing vector area for the near face is - dAn . The contribution to the above divergence integral from "far face n" is, approximately, [dAn]B(x) ≈ [ en(xfar) ( Πi≠n dx'i) ]B(xfar) = ( Πi≠n dx'i) en(xfar) B(xfar) where xfar is taken to be a point at the center of far face n. Recall from Section 7 (s) that B(x) = B'n(x')en where B'n(x') = en(x) B(x) which says that, when B is expanded on the en, the coefficients B'n of the expansion are the contravariant components of vector B transformed into B' in x'-space (the curvilinear coordinate space) by B' = RB. Inserting the last equation above applied at x' = x'far , B'n(x'far) = en(xfar) B(xfar) into the previous ≈ equation then gives dAnB(xfar) ≈ ( Πi≠n dx'i) B'n(x'far) This far face n contribution to the flux integral is now expressed entirely in terms of x'-space objects and coordinates. A similar expression obtains for the near face n, but the sign of dAn is reversed. Adding the contributions of these two faces of pair n gives ∫two faces n dAB(x) = { B'n(x'far) – B'n(x'near) } ( Πi≠n dx'i) In x'-space, if x'cen is a point at the center of the near face of face pair n, then x'near = x'cen x'far = x'cen + e'n dx'n where e 'n = axis-aligned basis vector in x'-space , since these two points map into the near and far face n centers in x-space. For any function f, f(x'far) - f(x'near) = (∂'n f(x')) dx'n // no implied sum on n where a change is made only in coordinate x'n by amount dx'n. Applying to f = J B'n yields { B'n(x'far) – B'n(x'near) } ≈ ∂'n [ B'n(x'cen)] dx'n In the limit that differentials are very close to 0, replace xcen by x. Then ∫two faces n dAB(x) = ∂'n [ B'n(x')] dx'n ( Πi≠n dx'i) = ∂'n [B'n(x')] ( Πi dx'i) where now all N differentials are present in ( Πi dx'i). The total flux flowing out through all N pairs of faces of the N-piped in x-space is this same result with an implicit sum on n, so total flux = ∫ dAB(x) = ∂'n [ B'n(x')] ( Πi dx'i) = ∂'n [B'n(x')] dV' where dV' = Πi dx'i is the volume of the differential N-piped in Cartesian-view x'-space (Section 8 (a)) . The divergence of B from the defining symbolic expression is then [div B](x) = ∫ dAB(x) / dV = ∂'n [ B'n(x')] (dV'/dV) where dV is the volume of the N-piped shown above. In Section 8 (h) it is shown that dV = (g')1/2 dV' => (dV'/dV) = 1/ so that [div B](x) = [1/] ∂'n [ B'n(x')] // all x'-space coordinates and objects [div B](x) = ∂nBn(x) // all x-space coordinates and objects The added second line just shows [div B](x) expressed in terms of the Cartesian x-space coordinates and objects, while the first line resulting from our derivation shows the same [div B](x) expressed in terms of only x'-space objects and coordinates. The claim advertised above has been fulfilled. If B is a tensorial vector, then as noted above div B is a tensorial scalar, [div B](x) = [div B]'(x') and thus the left sides of both equations above could be replaced by [div B]'(x'). (b) Various expressions for div B It is shown above that [div B](x) = [1/] ∂'n [ B'n(x')] To obtain div B written in terms of covariant components B'n, one sets B'n = g'nm B'm to get [div B](x) = [1/] ∂'n [g'nm(x') B'm(x')] and recall that the B'm are the coefficients of B when expanded on the en, B = B'nen. In practical work B is expanded on the unit vectors n ≡ en/ |en | = en/h'n so that B = B'nen = B'n h'nn = B'nn where B'n ≡ B'n h'n and then [div B](x) = [1/] ∂'n [B'n(x')/ h'n(x') ] For example, spherical coordinate work might use 1, 1, 1 = , , . As noted earlier, the components Bn(x) are not contravariant vector components since they don't quite transform properly: B'n = RnmBm (B'n / h'n) = Rnm (Bn / 1) B'n = h'n (Rnm Bn) Our Picture B results so far are these, assuming B is a tensorial vector, General: B'n(x') = RnmBm(x) x = F-1(x') ≡ x(x') x' = F(x) ≡ x'(x) [div B](x) = [1/] ∂'n [ B'n(x')] B = B'nen [div B](x) = [1/] ∂'n [B'n(x')/ h'n(x') ] B = B'n n [div B](x) = [1/] ∂'n [g'nm(x') B'm(x')] B = B'nen [div B](x) = ∂n Bn(x) // Bn = Cartesian components of B B = Bn [div B](x) = [div B]'(x') For orthogonal curvilinear coordinates, one has g'ij = h'i2 δi,j g'ij = h'i-2 δi,j det(g'ij) = Πih'i2 = (Πih'i) = h'1h'2....h'N so the above expressions can be written (the arguments x' of the h'n are now suppressed) Orthogonal: [div B](x) = [1/(Πih'i)] ∂'n [(Πih'i) B'n(x')] B = B'nen [div B](x) = [1/(Πih'i)] ∂'n [(Πih'i) B'n(x')/ h'n ] B = B'n n [div B](x) = [1/(Πih'i)] ∂'n [(Πih'i) B'n(x')/h'n2] B = B'nen [div B](x) = ∂n Bn(x) // Bn = Cartesian components of B B = Bn [div B](x) = [div B]'(x') Comment: Are the equations of the above "General" block valid if B is not a tensorial vector? For such a B one might try to make it be tensorial "by definition" as discussed in Section 2 (h). One would then go ahead and define B'n(x') ≡ RnmBm(x) and claim success. If such a definition does not result in an inconsistency, then such a B has been moved into the class of tensorial vectors. Example 1 of Section 2 (h) shows have such an inconsistency might arise, and it is interesting to see how that plays out here. Suppose F is non-linear so that R and g' = RRT are functions of x' and are not constants. Take B(x) = x (the identity field) and try to make it be contravariant by definition, x'n ≡ Rnmxm. The Cartesian divergence is then div B = ∂nBn(x) = ∂nxn = 3. But the first equation of the General block says div B = [1/] ∂'n [ x'n] = ∂'n x'n + [1/] x'n∂'n = 3 + other stuff and thus the two calculations for div B disagree. As noted earlier, x' ≡ Rx conflicts with x' = F(x) in the case of non-linear F. This comment can be applied to the tensorial character of the differential operators treated in later Sections. (c) Translation from Picture B to Picture M&S Picture M&S reflects the notation used by Moon & Spencer. In order to avoid a symbol conflict with the Cartesian tensor components, the Curvilinear (now u-space) components are displayed in italics. The rules for translation are replace x' by u everywhere replace ∂'n by ∂n meaning ∂/∂un ( exception: on a "Cartesian" line ∂n means ∂/∂xn) replace g' by g (both the scalar and the tensor) and hn' by hn put all primed tensor components (scalar, vector, etc) into unprimed italics (eg, B'n → Bn , f' → f ) After this translation, all unprimed tensor components are functions of x, while all italicized tensor components are functions of u. Here then are the translations of the two blocks above: (implied summation everywhere) General: now Bn(u) = RnmBm(x) x = F-1(u) ≡ x(u) u = F(x) ≡ u(x) [div B](x) = [1/] ∂n [ Bn] B = Bnen [div B](x) = [1/] ∂n [ Bn/ hn ] B = Bnn // M&S 1.06 [div B](x) = [1/] ∂n [gnm Bm] B = Bnen [div B](x) = ∂nBn // Bn = Cartesian components of B B = Bn = Bn [div B](x) = [div B](u) // transformation (scalar) Orthogonal: [div B](x) = [1/(Πihi)] ∂n [(Πihi) Bn] B = Bnen [div B](x) = [1/(Πihi)] ∂n [(Πihi) Bn / hn] B = Bn n [div B](x) = [1/(Πihi)] ∂n [(Πihi) Bn / hn2] B = Bnen Notice that the scalar function [div B]'(x') → [div B](u) according to the fourth rule above, and that the arguments of all hk(u) are suppressed. As an example, for N=3 the second line above becomes [div B](x) = [1/(h1h2h3)] { ∂1[h2h3 B1(u) ] + cyclic } B = Bn n where + cyclic means two other terms with 1,2,3 cyclically permuted. With the replacements B → E , Bn→ En hn → the equation marked above agrees with Moon & Spencer p 2 (1.06). Comment: The AaSansOutline font Bn used above for components of vectors expanded onto n is a bit clumsy and does not reproduce well in PDF files. It had to be something in upper case distinct from Bn and Bn. In practice one can replace Bn with a different symbol and then Bn is just a formal notation appearing in formulas. For example, in spherical coordinates (1,2,3) = (r,θ,φ) one can make the replacements B1, B2, B3 → Br, Bθ, Bφ and these then do not conflict with Bx, By, Bz or Br, Bθ, Bφ . See Section 14 Example 1 for another example. (d) Comparison of various authors' notations Different authors use different symbols for curvilinear coordinates. They usually use x-space as the Cartesian space, and then something like u-space or ξ-space as the curvilinear space: Curvilinear coords Cartesian coords Curvilinear space Picture C xn x(0)n x-space hn Picture B x'n xn x'-space h'n Moon & Spencer (M&S) p 2 un xn u-space Morse & Feshbach p 115 ξn xn ξ-space hn Margenau & Murphy p 192 qn xn q-space Qn These authors don't use any special notation to distinguish Cartesian from curvilinear components, nor is it always clear whether a component is a coefficident of a unit vector or not, so one must be careful. For example, on page 115 Morse & Feshbach simply say which compare to the above [div A](x) = [div A](u) = [1/(h1h2h3)] ∂n [h1h2h3 An(u) / hn] A = An n so presumably one should identify the M&F An with An, the coefficient of n. (e) The Christoffel derivation of div B This derivation is done in Picture C where the curvilinear coordinates are x so that ∂a means ∂/∂xa and g ≠ 1 is the curvilinear metric tensor, Mention of x(0)-space is completely avoided. In fact, even the x coordinate symbol rarely appears and could also have been removed. Function div B is in x-space, and once seen as a tensorial scalar, can be regarded as also being in x(0) space, [div B](0)(x(0)) = [div B] (x) // = Ba;a We start with the covariant derivative of a contravariant vector field component, which object is known to be a mixed rank-2 tensor (seeAppendix G (e) form 2, and also (g)), Bb;a ≡ ∂aBb + ΓbanBn where Γcab = ½ gcd( ∂agbd + ∂bgad – ∂dgab ) . Then the tensorial scalar object div B is defined by index contraction to be div B ≡ Ba;a ≡ ∂aBa + Γaan Bn Evaluation of Γaan gives (see Appendix G (g) ) Γaan = ½ gad( ∂agnd + ∂ngad – ∂dgan ) = ½ gad ∂ngad = ½ (1/g)∂ng = (1/) ∂n() Therefore div B = Ba;a ≡ ∂aBa + Γaan Bn = ∂nBn + (1/) ∂n()Bn = [1/] ∂n [ Bn] in agreement with the geometric derivation. Thus the entire distinction from ∂nBn is the second term above which arises from the affine connection Γban part of the covariant derivative. 10. The Gradient in curvilinear coordinates This Section is considered in the Picture B context, (a) Expressions for grad f The gradient of f is defined in Cartesian space by [grad f]n ≡ Gn ≡ ∂nf(x) G = grad f = ∂nf(x) Assuming f is a tensorial scalar field under F, then Gn = ∂nf(x) are covariant vector field components under F. Since the equation above is then a tensor equation, we know it is covariant in the sense of Section 7 (u). Therefore in x'-space it becomes [grad f]'n = G'n ≡ ∂'nf'(x') According to Section 7 (s) ( based on Section 6 (f)) vector G can be expanded as G = G'1e1 + G'2 e2 +... = ΣnG'n en where en G = G 'n where the coefficients G'n are the covariant components of G' in x'-space, which is the transformed G. One can thus write G(x) = [grad f](x) = G'n en = ∂'nf'(x') en = en∂'nf'(x') ≡ 'CL f'(x') 'CL ≡ en∂'n G(x) = [grad f](x) = ∂nf(x) = ∂nf(x) = f(x) ≡ ∂n where the second line shows the usual Cartesian form of the gradient. The first line shows how one could define a "curvilinear gradient" operator 'CL ≡ en∂'n, but this does not seem particular useful. To restate, [grad f](x) = ∂'nf'(x') en // Curvilinear [grad f](x) = ∂nf(x) // Cartesian In the first line, the reciprocal vectors en exist in x-space, but the coefficients are expressed entirely in terms of x'-space coordinates and objects. Since f(x) is a scalar field, f(x) = f'(x'), one could regard the derivative appearing in the first line as ∂'nf'(x') = ∂'nf(x(x')) where x(x') = F-1(x') The contravariant components of grad f are then easily obtained as G'i(x') = [grad f]'i(x') = ∂'if '(x') = g'ij(x') ∂'jf '(x') G = grad f = G'i(x') ei If an expansion on unit vectors is desired, the right equation on the last line can be written, G = [grad f](x) = G'iei = (G'i h'i) i = G'i i where G'i ≡ h'i G'i so then G'i(x')/h'i = [grad f]'i(x') = ∂'if '(x') = g'ij(x') ∂'jf '(x') G = grad f = G'i i Gathering up these results one gets G'i(x') = [grad f]'i(x') = ∂'if '(x') G = grad f = G'i(x') ei G'i(x') = [grad f]'i(x') = ∂'if '(x') = g'ij(x') ∂'jf '(x') G = grad f = G'i(x') ei G i(x') = h'i[grad f]'i(x') = h'i ∂'if '(x') = h'i G'ij(x') ∂'jf '(x') G = grad f = G'i(x') i which can be rewritten [grad f](x) = (∂'if '(x')) ei [grad f](x) = (∂'if '(x')) ei = g'ij(x') (∂'jf '(x')) ei [grad f](x) = h'i (∂'if '(x')) i = h'i g'ij(x') (∂'jf '(x')) i [grad f](x) = (∂if(x)) = f(x) // Cartesian [grad f]'i(x') = Rij[grad f]j(x) // transformation f '(x') = f(x) Again, in each of the first three forms above, G = [grad f](x) is being expressed as a linear combination of e vectors which are in x-space, but the coefficients are given entirely in terms of x'-space coordinates and objects. Since div B was a scalar quantity, this mixture of vectors in x-space with components in x'-space did not arise. For the orthogonal case g'ij = (1/h'i2) δi,j so the block above becomes [grad f](x) = (∂'if '(x')) ei [grad f](x) = (∂'if '(x')) ei = (1/h'i2) (∂'if '(x')) ei [grad f](x) = h'i (∂'if '(x')) i = (1/h'i) (∂'if '(x')) i [grad f](x) = (∂if(x)) // Cartesian [grad f]'i(x') = Rij[grad f]j(x) // transformation f '(x') = f(x) This above equations can be converted from Picture B to Picture M&S using the same rules given in Section 9 (c), which we repeat below replace x' by u everywhere replace ∂'n by ∂n meaning ∂/∂un ( exception: on a "Cartesian" line ∂n means ∂/∂xn) replace g' by g (both the scalar and the tensor) and hn' by hn put all primed tensor components (scalar, vector, etc) into unprimed italics (eg, B'n → Bn , f' → f ) After this translation, all unprimed tensor components are functions of x, while all italicized tensor components are functions of u. The translated results are then (all implied sums) [grad f](x) = (∂if ) ei [grad f](x) = (∂if ) ei = gij (∂jf ) ei [grad f](x) = hi(∂if ) i = hi gij (∂jf ) i [grad f](x) = (∂if) // Cartesian [grad f ]i(u) = Rij[grad f]j(x) // transformation (vector) f (u) = f(x) = f(x(u)) Notice that [grad f]'i(x') → [grad f]i(u) according to the fourth rule, meaning G'i(x') → Gi(u) . For orthogonal curvilinear coordinates, [grad f](x) = (∂if ) ei [grad f](x) = (∂if ) ei = (1/hi2) (∂if ) ei [grad f](x) = hi(∂if ) i = (1/hi) (∂if ) // M&S 1.05 One can always make the replacement f (u) = f(x(u)) in any of the above equations (f scalar). And one more time: the various e vectors are in x-space, but all the coefficients are expressed in curvilinear u-space coordinates and components. With the replacements f → φ i→ ai hi → the equation marked agrees with Moon & Spencer p 2 (1.05). (b) Expressions for grad f B Sometimes one is interested in the following quantity (back to Picture B) grad f B where f is a tensorial scalar field and B is a tensorial vector field. In this case, since grad f is a tensorial vector field, the quantity grad f B is a tensorial scalar, and so (grad f)' B' = (grad f) B . This quantity grad f B can be written several ways depending on how B is expanded: B = Σi B'i ei => grad f B = (∂'nf') en Σi B'i ei = (∂'nf') B'n B = Σi B'i ei => grad f B = (∂'nf') en Σi B'i ei = (∂'nf') B'n The last line can be written, using B'i ≡ B'i h'i , B = Σi [B'i h'i] i ≡ Σi B'i i => grad f B = (∂'nf') B'n = (∂'nf') (B'n/h'n) To summarize: [grad f](x) B(x) = (∂'n f'(x')) B'n(x') for B = Σi B'i ei [grad f](x) B(x) = (∂'n f'(x')) B'n(x') for B = Σi B'i ei [grad f](x) B(x) = (∂'n f'(x')) B'n(x')/h'n(x') for B = Σi B'i i [grad f](x) B(x) = (∂nf(x)) Bn(x) for B = Bn // Cartesian [grad f](x) B(x) = [grad f]'(x') B'(x') Since grad f B is a scalar, these results resemble the divergence results more than the gradient ones. Everything on the right side of the first three equations involves only x'-space coordinates and components. The conversion from Picture B to Picture M&F is straightforward (see Sections 8 and 9) [grad f](x) B(x) = (∂nf ) Bn B = Σi Bi ei [grad f](x) B(x) = (∂nf ) Bn B = Σi Bi ei [grad f](x) B(x) = (∂nf ) Bn/hn B = Σi Bi i [grad f](x) B(x) = (∂nf) Bn // Cartesian B = Bn [grad f](x) B(x) = [grad f ](u) B(u) // transformation (scalar) where once again f (u) = f(x(u)). Comment: According to the second line above, one can write grad f dx = ∂nf(x) dxn = df = f(x+dx) - f(x) This equation df = [grad f](x) dx is sometimes used as an alternate definition of the gradient. If dx is selected to be in the direction of grad f, the dot product has its maximum value, and therefore the gradient points in the direction of the maximum change of a scalar function f(x). For N=2, in the usual 3D plot of real f(x,y), the gradient then points "uphill", and the negative of the gradient then points "downhill". 11. The Laplacian in curvilinear coordinates The Laplacian (also known as the Laplace-Beltrami operator) is defined by lap f = div (grad f) = div G where G ≡ grad f Since f is (by assumption) a tensorial scalar field, grad f is a tensorial vector. Then, as found in Section 9, div(grad f) is a tensorial scalar, meaning [lap f](x) = [lap f]'(x'). In Cartesian coordinates one writes lap f = 2f = f = Σn∂n2f but this form gets modified when lap f is expressed in curvilinear coordinates. Section 9 showed that div G = [1/] ∂'m [G' m] where G = G'nen g' = det(g') Section 10 showed that grad f = G = [g'nm (∂'nf ') ] em = G' mem G' m = g'nm (∂'nf ') ∂'nf' = ∂'n f(x(x')) = ∂'n f '(x') Therefore lap f = div (grad f) = div G = [1/] ∂'m [G' m] = [1/] ∂'m [g'nm (∂'nf ')] so the general results can be concisely stated: [lap f](x) = [1/] ∂'m [ g'nm(x') (∂'nf '(x')) ] // implied sum on n and m [lap f](x) = Σn ∂n2f(x) // Cartesian f'(x') = f(x) = f(x(x')) [lap f](x) = [lap f]'(x') For an orthogonal coordinate system, g'nm = h'n2 δn,m g'nm = (1/h'n2) δn,m = Πi h'i and the first line above simplifies to [lap f](x) = [1/(Πih'i)] ∂'m [(Πih'i) (1/h'm2) (∂'mf ') ] Converting from Picture B to Picture MS gives (see Section 9 (c)) [lap f](x) = [1/] ∂m [gnm (∂nf ) ] [lap f](x) = Σn ∂n2f(x) // Cartesian f (u) = f(x) = f(x(u)) [lap f](x) = [lap f ](u) // transformation (scalar) The first line simplifies in the orthgonal case to [lap f](x) = [1/(Πihi)] ∂m [ (Πihi) (1/hm2) (∂mf ) ] // orthogonal // M&S 1.09 For N=3 this says, [lap f](x) = 1/(h1h2h3) { ∂1 [ (h2h3/h1) ∂1f ] + cyclic } With the replacements f → φ hi2 → gii (Πihi) → g the equation marked above agrees with Moon & Spencer p 3 (1.09). 12. The Curl in curvilinear coordinates The vector curl is defined only in N=3 dimensions (but see section (f) below). Picture B is used. In Cartesian coordinates one writes [curl B]i(x) = [ x B(x)]i = εijk∂jBk(x) but when expressed in terms of curvilinear coordinates and components, the form is different. (a) Definition of curl B Consider the x-space differential 3-piped shown on the right side of the figure in Section 8 (a), This 3-piped has three pairs of parallel faces. Within each pair, the "near" face touches the point x at which the tails of the spanning vectors meet, while the "far" face does not. As shown in Section 8 (c), the vector area associated with the face pair n is ( |J| = g'1/2 when g = 1) dAn = | det(Sij)| en ( Πi≠n dx'i) = |J| en ( Πi≠n dx'i) = g'1/2 en ( Πi≠n dx'i) n = 1,2,3 where J = det(Sij) is the Jacobian, as in Section 5 (k), and en is a reciprocal base vector, as in Section 6 or Appendix A (called En). Area dAn is the out-facing vector area for the far face of pair n, while - dAn is the out-facing vector for the near face. Consider now the line integral of a vector field B(x) around the boundary of near face n, where the circulation sense of the integral is determined from the right-hand-rule by the direction of dAn which is the same as the direction of en. For example, for the bottom face (near face 3) of the 3-piped shown above, this vector points "up", or toward the center of the 3-piped. Denote this line integral by ( Bdx)n Sometimes this line integral is referred to as "the circulation" or "the rotation" of B around near face n (and rot B is another notation used for curl B). In x-space the quantity C(x) ≡ curl B(x) is a vector field defined in the following manner in the limit that all the differentials dx'n → 0 : C dAn = ( Bdx)n C ≡ curl B As shown in Appendix D, C = curl B is in fact a vector field density of weight -1. Since dAn is given above in terms of en, C should be expanded on the ek and one writes (as in Appendix D ****), C = J-1 Σk=13 C'k ek so that C dAn = J-1 (Σk C'k ek) (g'1/2 en ( Πi≠n dx'i) ) = J-1C'n g'1/2 ( Πi≠n dx'i) and then J-1C'n(x') ( Πi≠n dx'i) = ( Bdx)n Our task is to compute this line integral and thereby come up with an expression for C'n(x'), the components of curl B when B is expanded onto the ek in x-space. (b) Computation of the line integral This shall be done for the bottom face (n=3) of the x-space differential 3-piped. Since e3 is "up", the circulation is a counterclockwise line integral around the boundary of the bottom face. It is useful to have the above picture near at hand to allow visualization of the four contributions to the line integral: ( Bds)3 ≈ [B(xfront) - B(xback)] (e1 dx'1) + [B(xright) - B(xleft)] (e2 dx'2) where B(xfront) refers to the value of B at the center of the "front" edge of the parallelogram which is the bottom face, and similarly for the other three edges. In the limit that the dx'n → 0, this simple approximation of the line integral is "good enough" to produce the desired results. Motivated by ei ej = δij, expand B as follows B = B'jej where B'j(x') = B(x) ej where the B'j are the covariant components of B in x'-space. This gives ( Bdx)3 = [B'1(x'front) - B'1(x'back)] dx'1 + [B'2(x'right) - B'2(x'left)] dx'2 where x'front = F(xfront) and similarly for the other three points. In x-space one has dxBF ≡ xback - xfront = e2dx'2 dxRL ≡ xright - xleft = e1dx'1 Applying matrix R gives the corresponding x'-space equations (recall dx' = R dx and e'n = Ren) dx'BF ≡ x'back - x'front = e'2dx'2 e'2 = (0,1,0...) dx'RL ≡ x'right - x'left = e'1dx'1 e'1 = (1,0,0...) Using the fact that f(x'+dx') ≈ f(x') + Σn∂nf(x') dx'n one finds B'1(x'back) ≈ B'1(x'front) + (∂B'1/∂x'2) dx'2 dx' = e'2dx'2 B'2(x'right) ≈ B'2 (x'left)) + (∂B'2/∂x'1) dx'1 dx' = e'1dx'1 so the circulation integral is then ( Bdx)3 ≈ – (∂B'1/∂x'2) dx'2 dx'1 + (∂B'2/∂x'1) dx'1 dx'2 = [– (∂B'1/∂x'2) + (∂B'2/∂x'1)] dx'1 dx'2 = [– ∂'2B'1 + ∂'1B'2] dx'1 dx'2 = [∂'1B'2 – ∂'2B'1] dx'1 dx'2 = ε3ab ∂'aB'b ( Πi≠3 dx'i) Repeating this calculation for faces 1 and 2 produces cyclic results, and all three face line integrals can be summarized as (where equality holds in the limit dx'i → 0) ( Bdx)n = εnab ∂'aB'b ( Πi≠n dx'i) Appendix D discusses the tensor ε known as the Levi-Cevita ε tensor. In Cartesian space, the up and down position of the indices does not matter, as for any tensor. In non-Cartesian space up and down does matter, as with any tensor. The only fact needed here is that ε'abc... = εabc... where ε' is the tensor in x'-space, as shown in Appendix D (d). In Cartesian space one can regard εabc... = εabc... as a bookkeeping permutation tensor with the properties given in Section 7 (h). Installing the prime on ε, ( Bdx)n = ε'nab ∂'aB'b ( Πi≠n dx'i) and this integral is then given entirely in terms of x'-space coordinates and objects. (c) Solving for the curl The equation for the curl obtained at the end of section (a) was J-1C'n(x') ( Πi≠n dx'i) = ( Bdx)n Insert the section (b) result for ( Bdx)n to get J-1C'n(x') ( Πi≠n dx'i) = ε'nab ∂'aB'b ( Πi≠n dx'i) The differentials cancel out, so then take dx'i→ 0 and thus shrink the 3-piped around the point of interest x = F-1(x') so that J-1C'n = [(1/) ε'nab ∂'aB'b ] C = J-1C'n en B = B'n en curl B = C = [(1/) ε'nab ∂'aB'b ] en The comparison between the curvilinear and Cartesian expressed curls is this: [curl B](x) = [(1/) ε'nab ∂'aB'b(x') ] en = (1/) { [∂'1B'2 - ∂'2B'1] e3 + cyclic } [curl B](x) = εnab ∂aBb(x) = { [ ∂1B2 - ∂2B1 ] + cyclic } Comment: In the first line above one can replace [curl B](x) by [ x B](x) with the understanding that the LHS is the curl in Cartesian x-space and the RHS is expressing this LHS in terms of x'-space coordinates and objects. The RHS is certainly not equal to ' x B' = ε'nab ∂'aB'b(x') ' . It is to avoid this possible confusion that the curl is written out as the word curl, and the same comment applies to the other differential operators. (d) Various forms of the curl The first form is that just presented above, J-1C'n = [(1/) ε'nab ∂'aB'b ] C = J-1C'n en B = B'n en curl B = [(1/) ε'nab ∂'aB'b ] en If it is desired to have contravariant components of B, one gets J-1C'n = [(1/) ε'nab ∂'a(g'bcB'c )] C = J-1C'n en B = B'n en curl B = [(1/) ε'nab ∂'a(g'bcB'c )] en For practical applications, one usually wants both vectors expanded on the n unit vectors in this way C = J-1C'n en = (J-1C'n h'n) n ≡ C 'n n C 'n = h'n J-1C'n B = B'n en = (B'n h'n) n ≡ B'n n B'n = h'nB'n so that C'n = [(1/) h'n ε'nab ∂'a(g'bc B'c/h'c )] C = C 'n n B = B 'n n curl B = [(1/) h'n ε'nab ∂'a(g'bc B'c/h'c )] n curl B = C To summarize: B'c = RcdBd g' = g'(x') etc. [curl B](x) = ε'nab [(1/) ∂'aB'b ] en B = B'nen [curl B](x) = ε'nab [(1/) ∂'a(g'bcB'c )] en B = B'nen [curl B](x) = ε'nab [(1/) h'n ∂'a(g'bc B'c/h'c )] n B = B 'n n [curl B](x) = εnab ∂aBb(x) // Cartesian B = Bn Converting from Picture B to Picture MS (see Section 9 (c)) one gets : [curl B](x) = εnab [(1/) ∂aBb ] en B = Bnen [curl B](x) = εnab [(1/) ∂a(gbcBc )] en B = Bnen [curl B](x) = εnab[(1/) hn ∂a(gbc Bc/hc )] n B = Bnn [curl B](x) = εnab ∂aBb(x) // Cartesian B = Bn Warning: The object εnab in the first three equations is now in u-space which is non-Cartesian, so up and down index positions do matter, but when indices are all up, it continues to be the normal permutation tensor. Each of the above results can be written as a determinant using the idea det(Q) ≡ Σi Q1i cof(Q1i) : [curl B](x) = (1/) B = Bn en [curl B](x) = (1/) B = Bn en [curl B](x) = (1/) B = Bn n // M&S 1.07 [curl B](x) = // here ∂n= ∂/∂xn and Bi = Bi(x) B = Bn (e) The curl in orthogonal coordinate systems For such systems gij = δi,jhi2 and det(gab) = h12h22h32 so = h1h2h3 . It is then a simple matter to convert all the above forms and the results are: Picture B: B'c(x') = RcdBd(x) hi' = hi'(x') etc. [curl B](x) = ε'nab [(1/(h1'h2'h3') ∂'aB'b ] en B = B'nen [curl B](x) = ε'nab [(1/(h1'h2'h3')) ∂'a(h'b2B'b )] en B = B'nen [curl B](x) = ε'nab [(1/(h1'h2'h3')) h'n ∂'a(h'b B'b) )] n B = B 'n n [curl B](x) = εnab ∂aBb(x) // Cartesian B = Bn Picture M&S: Bc(u) = RcdBd(x) hi = hi(u) etc. [curl B](x) = εnab [(h1h2h3)-1 ∂aBb ] en B = Bnen [curl B](x) = εnab [(h1h2h3)-1 ∂a(hb2Bb )] en B = Bnen [curl B](x) = εnab[(h1h2h3)-1 hn ∂a(hb Bb)] n B = Bnn [curl B](x) = (h1h2h3)-1 B = Bn en [curl B](x) = (h1h2h3)-1 B = Bn en [curl B](x) = (h1h2h3)-1 B = Bn n // M&S 1.07a With the replacements B → E Bn→ En n→ an hi → (h1h2h3)-1 → (1/) the equations marked above agree with Moon & Spencer p 2 (1.07) and p 3 (1.07a). (f) The curl in N > 3 dimensions Looking at the basic form of the curl above [curl B]n(x) = εnab ∂aBb(x) // Cartesian it is hard to imagine a generalization to N>3 dimensions where the curl is still a vector. The only vectors available for construction purposes are ∂n and Bn . For N=4 one might try out various generalizing forms [curl B]n(x) = (1/) εnabc ∂a(∂b Bc) = (1/) εnabc ∂a∂b Bc ? [curl B]n(x) = (1/) εnabc ∂a (BbBc)) ? but these two forms vanish because antisymmetric ε is contracted against something symmetric. Thus the idea of using multiple cross products as used in Appendix A does not prove helpful. The rank-2 tensor Bb;a – Ba;b = ∂aBb – ∂bBa discussed in Appendix D (h) provides the logical extension of the curl to N > 3 dimensions. For N=3 it happens that the object can be associated with a vector, [curl B]n = εnab [Bb;a – Ba;b ]/2 = εnab Bb;a = εnab [∂aBb – ∂bBa ]/2 = εnab∂aBb . In relativity work, since N=4, there is no vector curl, and one sees Bb;a – Ba;b referred to as the covariant curl, and ∂aBb – ∂bBa as the ordinary curl ( Weinberg p 106). Writing the N-dimensional contravariant curl components in this manner in Cartesian x-space, [curl B]ij = (Bj;i – Bi;j) one can then ask how this generalized curl would be expressed in terms of x'-space coordinates and objects. (This curl is a regular rank-2 tensor with weight 0, the vector curl had weight -1 ). The general issue of expanding tensors is addressed in Appendix E where this general result is obtained A = Σijk... A'ijk... (eiejek...) A'ijk... = contravariant components of A in x'-space Applying this to A = [curl B] one gets [curl B] = Σij[curl B]'ijeiej = Σij(B'j;i – B'i;j)eiej The Cartesian components of curl B in x-space can then be expressed in terms of x'-space components and coordinates, [curl B]ab(x) = Σij [curl B]'ij (ei)a(ej)b = Σij(B'j;i – B'i;j)(ei)a(ej)b where the tangent base vectors en exist as usual in x-space, and B'i = B'i(x') where x' = F(x). 13. The Vector Laplacian in curvilinear coordinates This operator is defined in terms of the vector curl which is only defined for N=3. The context is Picture B. (a) Derivation of the Vector Laplacian in general curvilinear coordinates The definition of the vector Laplacian of a vector field B(x) is 2B ≡ grad(div B) – curl (curl B) , so, as expected, the vector Laplacian is a vector field. In Cartesian coordinates, one finds that [2B]i = 2(Bi) ≡ Σn ∂n2Bi but expressed in general curvilinear coordinates the form gets modified. To avoid confusion, some authors use different symbols for the vector Laplacian operator. For example, M&S use in place of 2 and we will honor these authors by using that symbol here, so B ≡ grad(div B) – curl (curl B) In order to make use of the results of earlier sections, define G ≡ grad(f) where f = div B V ≡ curl C where C ≡ curl B so that B = G – V Section 10 (c) gives this expression for G, G = grad(f) = (∂'kf ') ek in which expression Section 9 (b) allows replacement of f ' as follows, f ' = f '(x') = f(x) = div B = [1/] ∂'i [ B'i] so that => G = ∂'k{ (1/) ∂'i ( B'i)} ek // 2 implied sums, i and k The second term V is little more complicated. First, from Section 12 (d), C = curl B = ε'nab [(1/) ∂'a{B'b} ] en = C'n en V = curl C = ε'ncd [(1/) ∂'c{C'd} ] en = V'n en where recall from Appendix D (d) that ε'abc... = εabc.. = εabc.. = the usual permutation tensor, but written up and primed so as to be in covariant form. Appendix D (h) shows that C and V are both true contravariant vectors, with components C'n and V'n in x'-space. The first line above says C'e = ε'eab [(1/) ∂'aB'b ] so that C'd = g'de C'e = g'de ε'eab [(1/) ∂'aB'b ], which can then be inserted then into the second line to get V = curl C = ε'ncd [ (1/) ∂'c{ g'de ε'eab [(1/) ∂'aB'b ]} ] en = V'n en = (1/) ε'ncd ε'eab ∂'c{ (1/) g'de (∂'aB'b) } en // note g'dc = g'dc(x') , etc which has 6 implied sums, a,b,c,d,e, and n. For each value of n, there are not really 35 terms because most terms vanish due to the ε factors. Looking at ε'ncd ε'eab = εncd εeab, one sees that for each n, the c and d sums generate only 2 terms, and for each of these εeab generates 3*2*1 = 6 terms, so there are 12 terms total for each n. Later when g'de is assumed diagonal, the effective factor is εncd εeab implying 2 * (2*1) = 4 terms, which shall be written out in that case. Combining these terms, the vector Laplacian is now B = G – V = ∂'k{ (1/) ∂'i ( B'i)} ek – (1/) ε'ncd ε'eab ∂'c{ (1/) g'de (∂'aB'b) } en Setting ek = g'kn en in the first term gives B = ∂'k{ (1/) ∂'i ( B'i)} g'kn en – (1/) ε'ncd ε'eab ∂'c{ (1/) g'de (∂'aB'b) } en = en [g'kn ∂'k{ (1/) ∂'i ( B'i)} – (1/) ε'ncd ε'eab ∂'c{ (1/) g'de (∂'aB'b) } ] so at least now both terms use the same expansion base vector en . As a next step, en = h'n n so B = h'n n [g'kn ∂'k{ (1/) ∂'i ( B'i)} – (1/) ε'ncd ε'eab ∂'c{ (1/) g'de (∂'aB'b) } ] The component B'b in the second term can be made contravariant using B'b = g'bfB'f to get B = h'n n [g'kn ∂'k{ (1/) ∂'i ( B'i)} – (1/) ε'ncd ε'eab ∂'c{ (1/) g'de (∂'a [g'bfB'f]) } ] and then, as was done in earlier sections, replace B'n = b'n/h'n B = B'n n to get this final form in "practical units", B = h'n n [g'kn ∂'k{ (1/) ∂'i (B'i/h'i)} – (1/) ε'ncd ε'eab ∂'c{ (1/) g'de (∂'a [g'bf B'f/h'f]) } ] There are so many options here it is difficult to summarize, but here are two forms from above: [ B](x) = en [g'kn ∂'k{ (1/) ∂'i ( B'i(x'))} – (1/) ε'ncd ε'eab ∂'c{ (1/) g'de (∂'a [g'bfB'f(x')]) } ] B = B'n en [ B](x) = n h'n [g'kn ∂'k{ (1/) ∂'i (B'i(x')/h'i)} – (1/) ε'ncd ε'eab ∂'c{ (1/) g'de (∂'a [g'bf B'f(x')/h'f]) } ] B = B'n n [ B](x) = 2(B(x)) // Cartesian, meaning [ B]i(x) = 2Bi(x) B = Bn Converting from Picture B to Picture MS gives (see Section 9 (c)) [ B](x) = en [gkn ∂k{ (1/) ∂i (Bi)} – (1/) εncd εeab ∂c{ (1/) gde (∂a [gbfBf]) } ] B = Bn en [ B](x) = n hn [gkn ∂k{ (1/) ∂i (Bi/hi)} – (1/) εncd εeab ∂c{ (1/) gde (∂a [gbf Bf/hf]) } ] B = Bn n [ B](x) = 2(B(x)) // Cartesian, meaning [ B]i(x) = 2[Bi(x) ] B = Bn In the first two equations above, all the ∂i mean ∂/∂ui and the argument u of all functions is suppressed. The Cartesian form will be verified below. (b) The Vector Laplacian in orthogonal curvilinear coordinates We continue in Picture M&A and process only the second equation of the above block, since it is the one with the practical components B = Bn n . Setting gij = hi2δi,j and gij = (1/hi)2δi,j things simplify somewhat B = hn n [gkn ∂k{ (1/) ∂i (Bi(u)/hi)} – (1/) εncd εeab ∂c{ (1/) gde (∂a [gbf Bf(u)/hf]) } ] = hn n [δk,n ∂k{ (1/) ∂i (Bi(u)/hi)} (1/hn)2 – (1/) εncd εeab ∂c{ (1/) hd2δd,e (∂a [hb2δb,f Bf(u)/hf]) } ] = hn n [ ∂n{ (1/) ∂i (Bi(u)/hi)} (1/hn)2 – (1/) εncd εeab ∂c{ (1/) hd2 (∂a [hb Bb(u)]) } ] = n [ (1/hn) ∂n{ (1/) ∂i (Bi/hi)} – (hn/) εncd εeab ∂c{ (1/) hd2 (∂a [hb Bb]) } ] The first term can be written as n (1/hn) ∂nT where T = (1/) ∂i (Bi/hi) To expand the second term, set n = 1 and then write things out explicitly. For the moment, we suppress the leading factor – (h1/) and write ε1cd εdab ∂c{ (1/) hd2 (∂a [hb Bb]) } = ε3ab ∂2{ (1/) h32 (∂a [hb Bb]) } – ε2ab ∂3{ (1/) h22 (∂a [hb Bb]) } c=2 d=3 c=3 d=2 = [ ∂2{ (1/) h32 (∂1 [h2 B2]) } – ∂2{ (1/) h32 (∂2 [h1 B1]) } ] a = 1 b = 2 a = 2 b = 1 – [ ∂3{ (1/) h22 (∂3 [h1 B1]) } – ∂3{ (1/) h22 (∂1 [h3 B3]) } ] a = 3 b = 1 a = 1 b = 3 = ∂2{ (1/) h32 ( ∂1 [h2 B2] – ∂2 [h1 B1] ) } – ∂3{ (1/) h22 (∂3 [h1 B1] – ∂1 [h3 B3] ) } Define now Γn in cyclic fashion. Γ1 ≡ (1/) h12 ( ∂2 [h3 B3] – ∂3 [h2 B2] ) Γ2 ≡ (1/) h22 ( ∂3 [h1 B1] – ∂1 [h3 B3] ) Γ3 ≡ (1/) h32 ( ∂1 [h2 B2] – ∂2 [h1 B1] ) and then we have shown that 2nd term (n=1) = – (h1/) ε1cd εdab ∂c{ (1/) hd2 (∂a [hb Bb]) } 1 = – (h1/) (∂2 Γ3 – ∂3 Γ2) 1 = + (h1/) (∂3 Γ2 – ∂2 Γ3) 1 Therefore the entire first term (n=1) of B is given by B (first term) = [(1/h1) ∂1T + (h1/) (∂3 Γ2 – ∂2 Γ3) ] 1 The other two terms are obtained by cyclic permutation so the final result is then [ B](x) = [(1/h1) ∂1T + (h1/) (∂3 Γ2 – ∂2 Γ3) ] 1 + cyclic = [(1/h1) ∂1T + (h1/) (∂3 Γ2 – ∂2 Γ3) ] 1 + [(1/h2) ∂2T + (h2/) (∂1 Γ3 – ∂3 Γ1) ] 2 + [(1/h3) ∂3T + (h3/) (∂2 Γ1 – ∂1 Γ2) ] 3 // M&S 1.11 where T = (1/) ∂i (Bi/hi) Γ1 = (1/) h12 ( ∂2 [h3 B3] – ∂3 [h2 B2] ) Γ2 = (1/) h22 ( ∂3 [h1 B1] – ∂1 [h3 B3] ) Γ3 = (1/) h32 ( ∂1 [h2 B2] – ∂2 [h1 B1] ) With the replacements B → E Bn→ En n→ an hi → T → ϒ the result agrees with M&S p 3 (1.11). A more compact summary is this: B = [ (1/hn) ∂nT – (hn/) εnab∂a Γb ] n T = (1/) ∂i (Bi(u)/hi) Γb = (1/) hb2 εbcd ( ∂c [hd Bd(u)]) (c) The Vector Laplacian in Cartesian coordinates First, one can verify that the last result of section (b) gives the starting point formula for B if g = 1 (in u-space). One then has, hi = 1 g = 1 u = F(x) = x (n)i = Sni = δni => n = , B = Bn n = Bn => Bn = Bn so the above 3-line equation block becomes B = [ ∂nT – εnab∂a Γb ] T = ∂i (Bi(u)) Γb = εbcd ( ∂c Bd(u)) or B = [ ∂n{∂iBi} – εnab∂a { εbcd ( ∂c Bd)} ] (*) = [ ∂n{div B} – εnab∂a { (curl B)b } ] = [ ∂n{div B} – [curl (curl B)]n } ] = {div B} – [curl (curl B)] QED Second, one can verify the claim made earlier that in Cartesian coordinates [ B]n = 2 Bn . To show this, it is necessary to show that (left side from (*) above) ∂n ∂i Bi – εnab∂a εbcd(∂c Bd) = ∂i2Bn ? εnab εbcd ∂a (∂c Bd) = ∂n ∂i Bi – ∂i2Bn ? εbna εbcd ∂a (∂c Bd) = ∂n ∂i Bi – ∂i2Bn ? But since index b now appears only in the ε's, use ( up and down indices same in Cartesian x-space) εbna εbcd = δncδad – δndδac // Appendix D (j) item 4 so (δncδad – δndδac) ∂a (∂c Bd) = ∂n ∂i Bi – ∂i2Bn ? δncδad∂a (∂c Bd) – δndδac∂a (∂c Bd) = ∂n ∂i Bi – ∂i2Bn ? ∂a (∂n Ba) – ∂a (∂a Bn) = ∂n ∂i Bi – ∂i2Bn ? ∂n (∂a Ba) – ∂a2Bn = ∂n (∂i Bi) – ∂i2Bn ? Since this last equation is true on inspection, QED. 14. Summary of Differential Operators in curvilinear coordinates The results are given in the Picture M&S context, and are copied from Sections 9-13. The Standard Notation of Section 7 is used throughout. In all the differential operator equations below, an operator acts either on a tensorial vector field B or on a tensorial scalar field f. On the right side of the drawing above, objects are said to be in x-space, and f(x) and Bn(x) (components of B) are x-space tensorial objects. On the left side of the drawing objects are said to be in u-space. The function f is represented in u-space as f (u), while there are three different ways to represent the components of B, called Bn(u), Bn(u) and Bn(u). There is a big distinction between the x-space objects and the u-space objects. For the scalar, f (u) = f(x) = f(x(u)) and so f has a different functional form than f. For the vector components, the u-space components are linear combinations of the x-space components, for example Bn = RnmBm ( or B = RB, contravariant vector transformation). See comment at the end of Section 9 (c) concerning the font used for Bn(u). On lines marked "Cartesian", ∂n = ∂/∂xn and f(x) and Bn(x) appear (Cartesian components). On other lines, ∂n = ∂/∂un, and the f and B objects appear in italics and are functions of u. The other functions like hn, gab and g are also functions of u. The vectors en, n and en all exist in Cartesian x-space. The en are the tangent base vectors of Section 3, and the en are the reciprocal base vectors of Section 6. The unit vectors n ≡ en/ |en| = en/hn are used as well. The dot product AB is the covariant one of Section 5 (i). For each differential operator, the object on the LHS of the equations is always the same: it is a differential operator acting on f(x) or B(x) in x-space. In the Cartesian lines, the RHS expresses that LHS object in terms of Cartesian objects and Cartesian coordinates. On the other lines, the RHS expresses that exact same LHS x-space object in terms of Curvilinear (u-space) objects and coordinates. When the LHS is a scalar, the LHS object can be considered to be in either x-space or u-space. When the LHS is a vector, that LHS object is in x-space but can be related to u-space objects by a linear transformation by R. The expressions marked below appear on pages 2 or 3 of Moon & Spencer (M&S). general: g ≡ det(gab) hn2 ≡ gnn ∂i = gij∂j Bn = hnBn en = hnn orthogonal: = (Πihi) = h1h2...hN gnm = hn2 δn,m gnm = hn-2 δn,m ___________________________________________________________________________ (a) divergence divergence general: [div B](x) = [1/] ∂n [ Bn] B = Bnen [div B](x) = [1/] ∂n [ Bn/ hn ] B = Bnn // M&S 1.06 [div B](x) = [1/] ∂n [gnm Bm] B = Bnen [div B](x) = ∂nBn // Bn = Cartesian components of B B = Bn = Bn [div B](x) = [div B](u) // transformation (scalar) divergence orthogonal: [div B](x) = [1/(Πihi)] ∂n [(Πihi) Bn] B = Bnen [div B](x) = [1/(Πihi)] ∂n [(Πihi) Bn / hn] B = Bn n [div B](x) = [1/(Πihi)] ∂n [(Πihi) Bn / hn2] B = Bnen divergence orthogonal N=3: [div B](x) = [1/(h1h2h3)] { ∂1[h2h3 B1(u) ] + cyclic } B = Bn n ___________________________________________________________________________ (b) gradient and gradient dot vector gradient general: [grad f](x) = (∂if ) ei [grad f](x) = (∂if ) ei = gij (∂jf ) ei [grad f](x) = hi(∂if ) i = hi gij (∂jf ) i [grad f](x) = (∂if) // Cartesian [grad f ]i(u) = Rij[grad f]j(x) // transformation (vector) f (u) = f(x) = f(x(u)) gradient orthogonal: [grad f](x) = (∂if ) ei [grad f](x) = (∂if ) ei = (1/hi2) (∂if ) ei [grad f](x) = hi(∂if ) i = (1/hi) (∂if ) // M&S 1.05 gradient dotted with a vector: [grad f](x) B(x) = (∂nf ) Bn B = Bnen [grad f](x) B(x) = (∂nf ) Bn B = Bnen [grad f](x) B(x) = (∂nf ) Bn/hn B = Bn n [grad f](x) B(x) = (∂nf) Bn // Cartesian B = Bn [grad f](x) B(x) = [grad f ](u) B(u) // transformation (scalar) ___________________________________________________________________________ (c) Laplacian Laplacian general: [lap f](x) = [1/] ∂m[gnm (∂nf ) ] [lap f](x) = ∂n2f(x) // Cartesian f (u) = f(x) = f(x(u)) [lap f](x) = [lap f ](u) // transformation (scalar) Laplacian orthogonal: [lap f](x) = [1/(Πihi)] ∂m[ (Πihi) (1/hm2) (∂mf ) ] // orthogonal // M&S 1.09 Laplacian orthogonal N=3: [lap f](x) = 1/(h1h2h3) { ∂1 [ (h2h3/h1) ∂1f ] + cyclic } ___________________________________________________________________________ (d) curl curl general: // N=3 only [curl B](x) = εnab [(1/) ∂aBb ] en B = Bnen [curl B](x) = εnab [(1/) ∂a(gbcBc )] en B = Bnen [curl B](x) = εnab[(1/) hn ∂a(gbc Bc/hc )] n B = Bnn [curl B](x) = εnab ∂aBb(x) // Cartesian B = Bn [curl B](x) = (1/) B = Bnen [curl B](x) = (1/) B = Bnen [curl B](x) = (1/) B = Bnn // M&S 1.07 [curl B](x) = // here ∂n= ∂/∂xn and Bi = Bi(x) B = Bn curl orthogonal: [curl B](x) = εnab [(h1h2h3)-1 ∂aBb ] en B = Bnen [curl B](x) = εnab [(h1h2h3)-1 ∂a(hb2Bb )] en B = Bnen [curl B](x) = εnab[(h1h2h3)-1 hn ∂a(hb Bb)] n B = Bnn [curl B](x) = (h1h2h3)-1 B = Bnen [curl B](x) = (h1h2h3)-1 B = Bnen [curl B](x) = (h1h2h3)-1 B = Bnn // M&S 1.07a ___________________________________________________________________________ (e) vector Laplacian vector Laplacian general: // N=3 only [ B](x) = en [gkn ∂k{ (1/) ∂i (Bi)} – (1/) εncd εeab ∂c{ (1/) gde (∂a [gbfBf]) } ] B = Bnen [ B](x) = n hn [gkn ∂k{ (1/) ∂i (Bi/hi)} – (1/) εncd εeab ∂c{ (1/) gde (∂a [gbf Bf/hf]) } ] B = Bnn [ B](x) = 2(B(x)) // Cartesian, meaning [ B]i(x) = 2[Bi(x) ] B = Bn vector Laplacian orthogonal: [ B](x) = [(1/h1) ∂1T + (h1/) (∂3 Γ2 – ∂2 Γ3) ] 1 // M&S 1.11 + [(1/h2) ∂2T + (h2/) (∂1 Γ3 – ∂3 Γ1) ] 2 +[(1/h3) ∂3T + (h3/) (∂2 Γ1 – ∂1 Γ2) ] 3 T = (1/) ∂i (Bi/hi) Γ1 = (1/) h12 ( ∂2 [h3 B3] – ∂3 [h2 B2] ) Γ2 = (1/) h22 ( ∂3 [h1 B1] – ∂1 [h3 B3] ) Γ3 = (1/) h32 ( ∂1 [h2 B2] – ∂2 [h1 B1] ) or [ B](x) = [ (1/hn) ∂nT – (hn/) εnab∂a Γb ] n T = (1/) ∂i (Bi/hi) Γb = (1/) hb2 εbcd ( ∂c [hd Bd]) ___________________________________________________________________________ Example 1: Polar coordinates: a practical curvilinear notation From earlier versions of this example we know that general: g ≡ det(gab) hn2 ≡ gnn ∂i = gij∂j Bn = hnBn en = hnn orthogonal: = (Πihi) = h1h2...hN gnm = hn2 δn,m gnm = hn-2 δn,m e1 = r(-sinθ,cosθ) = eθ = r θ // = r e2 = (cosθ,sinθ) = er = r // = θ r x y gij = R = u1 = θ u2 = r h1 = hθ = = r h2 = hr = = 1 Assume one is working with a 2D vector velocity field v(x), our first encounter with a "lower case" vector field which we have been careful to support with the general notations above. Since one knows the names of the variables 1= θ and 2= r, one might define the following new variables on the first line to be the officially named variables on the second line vθ vr vθ vr vθ vr vx vy v1 v2 v1 v2 v1 v2 v1=v1 v2=v2 contravariant covariant unit vector Cartesian (italic) (italic) (non-italic) (non-italic) One ends up with the comfortable v = vθ + vr notation as shown below. vθ ≡ v1 = h1ν1 = hθ vθ = r vθ // unit vector projection components vr ≡ v2 = h2ν2 = hr vr = vr vθ = Rθxvx + Rθyvy = -sinθ/r vx + cosθ/r vy vr = Rrxvx + Rryvy = cosθ vx +sinθ vy so vθ = r vθ = -sinθ vx + cosθ vy vr = vr = cosθ vx +sinθ vy v = vnen = vθeθ + vrer v = vnn = vθθ + vrr = vθ + vr v = vn = vx + vy As an example of a differential operator, consider the divergence for orthogonal coordinates from the above table, [div B](x) = [1/(Πihi)] ∂n [(Πihi) Bn / hn] B = Bn n which applied to the present situation reads (h1 = hθ = r and h2 = hr = 1), [div v](x) = [1/(h1h2)] { ∂1 [h2 v1] + ∂2[h1 v2] } V = vn n = [1/(hθhr)] { ∂θ [hr vθ ] + ∂r[hθ vr] } = (1/r) { ∂θvθ + ∂r(rvr) } Suppose vx and vy are constants. Then the Cartesian expression says [div v](x) = ∂nvn = ∂xvx + ∂yvy = 0 + 0 = 0 The above curvilinear expression gives [div v](x) = (1/r) { ∂θvθ + ∂r(rvr) } = (1/r) { ∂θ[-sinθ vx + cosθ vy] + ∂r(r [cosθ vx + sinθ vy]) } = (1/r) { [-cosθ vx - sinθ vy] + [cosθ vx + sinθ vy] } = 0 The notation illustrated here works for any curvilinear coordinates. Appendix A: Reciprocal Base Vectors the Hard Way Note: This Appendix is written in the development notation, not the Standard Notation, though a few equations are translated to the latter form. The rules for translation to Standard Notation are En → en Rij → Rij Sij → Sij 'nm → g'nm g'nm → g'nm . Introduction In Section 6 of the main text the reciprocal base vectors are defined as En ≡ g'ni ei , and the results given in that section, (en)k = Skn en em = 'nm |en| = = h'n S = [e1, e2, e3 .... eN ] (En)i ≡ gia Rna En Em = g'nm |En| = R = [1, 2, 3 .... N ]T = g'na Sia en Em = δn,m En ≡ g'ni ei en = 'ni Ei , are all applicable in the Picture A context with arbitrary metric tensors g' and g, This Appendix begins with a different definition of something called Ek. Although the definition is meaningful in the general Picture A context, the object so defined only agrees with the Ek of Section 6 if x-space is Cartesian (g = 1). The reason can be traced to the fact that the dot product rule en Em = δn,m is only valid for the Appendix A definition of Em when g = 1 because only then is a cross product orthogonal to all its component vectors. The main application of the reciprocal base vectors is in the study of curvilinear coordinates where one always takes g = 1, and g' is then the curvilinear coordinates metric tensor of interest. Therefore, the reader should think of this Appendix in the context of Picture B (a) Definition of En The reciprocal base vectors are defined in the following very strange looking and clumsy manner, (Ek)α ≡ det(R) (-1)k-1 εαiii...i...i (e1)i (e2)i ...... (ek)i.......... (eN)i where N is the number of dimensions of the Cartesian x-space RN in which the vectors en and En exist. Notice that the ε subscript ik is "crossed out" and the same for factor (ek)i . Crossed out means they are simply missing, they are omitted. Thus, in the above expression there are N-1 implied summation indices (α is fixed) and there are N-1 factors of the form (en)i . The object ε has N subscripts and is the "totally antisymmetric tensor" in N dimensions: ε123...N ≡ +1, and each time any two indices on ε are swapped, ε negates. For example, ε1234 = 1 but ε1432= -1. If two indices are the same, then ε = 0. (b) Simpler notation To avoid dealing with subscripts on subscripts, one can rewrite the above definition in a less precise but simpler notation (Ek)α ≡ det(R) (-1)k-1εαabc...x (e1)a(e2)b ...... (eN)x // κ(k) and (ek)κ are missing In this notation, subscript x stands for the Nth letter of the alphabet (imagine N ≤ 26). If κ is the kth letter of the alphabet, then κ is missing from the indices on ε, and the factor (ek)κ is missing from the product of factors. For example, if k = 2, then summation index κ = b is missing from the ε. Now take the ε subscript α and slide it right to the "hole" where κ is missing, picking up a minus sign for each step of this slide. Moving k-1 positions results in (-1)k-1. Thus the above becomes, (Ek)α ≡ det(R) εabc..α..x (e1)a(e2)b ...... (eN)x // (ek)κ is missing, α in κ position (the kth) Example: For N = 3 the above becomes, (E1)α ≡ det(R) εαbc(e2)b(e3)c => E1 = det(R) e2 x e3 a is missing (E2)α ≡ det(R) εaαc(e1)a(e3)c => E2 = det(R) e3 x e1 b is missing (E3)α ≡ det(R) εabα(e1)a(e2)b => E3 = det(R) e1 x e2 c is missing and the results are cyclic. Here is a detail from the middle line εaαc(e1)a(e3)c = – εαac(e1)a(e3)c = + εαca(e1)a(e3)c = εαca(e3)c (e1)a = [ e3 x e1]α (c) Generalized Cross Product of N-1 vectors of dimension N One can define a generalized "cross product" of N-1 vectors, each of dimension N, in this fashion: Qa ≡ εabc...x BbCcDd.....Xx where x and X represent the Nth letter of the alphabet. The ε object is again the totally antisymmetric tensor with N indices. In vector notation one writes this symbolically as Q = B x C x D x ... x X / N-1 factors, N-2 crosses This vector notation is defined by the previous line. The vector Q is orthogonal to all the vectors from which it is constructed! For example (here is the point where Q C ≡ gabQaCb needs to be QaCa, so g = 1 is required in x-space) Q C = CaQa = Ca εabc...x BbCcDd.....Xx = BbDd...Xx { εabc...x CaCc } But {..} is the contraction of something symmetric under a↔c (CaCc) with something antisymmetric under a↔c (εαabc...x) and therefore {..} = 0. In general SacAac = Sca Aca // relabel both dummy summation indices = Sac (-Aac) // S is Symmetric, A is antisymmetric = - Sac Aac // = the negative of the starting expression = 0 Similarly, QA = 0, QB = 0 and so on. Swapping the position of any two vectors in the generalized cross product causes Q to change sign. For example, swapping B and C, Qa ≡ εabc...x CbBcDd.....Xx = εacb...x CcBbDd.....Xx // b ↔ c = - εabc...x BbCcDd.....Xx = -Qa // swap indices on ε Thus, the notions of orthogonality and interchange are consistent with the regular Q = B x C cross product for N=3. When N=2, one must be a little careful with this notation. The component equation is Qa ≡ εab Bb => Q1 = B2 and Q2 = -B1 One might be tempted to express the vector equation as Q = B since there are no "no crosses". This vector equation is wrong, while the component equation is correct. One can rescue the vector notation by a simple trick. When N=2 the vector B can be represented of course as B = B1 + B2 . Imagine this 2D space to be embedded in the usual 3D space with a third axis . Then consider this 3D cross product: Q = B x => Qa ≡ εabc Bb()c = εabc Bbδ3,c = εab3 Bb = εabBb Thus, this trick reproduces the correct component equation, and it makes more obvious the fact that Q is orthogonal to B. Summary: The generalized cross product Q of N-1 vectors each of dimension N can be expressed in both component and vector notation: Qa ≡ εabc...x BbCcDd.....Xx Q = B x C x D x ... x X / N-1 factors, N-2 crosses Q is orthogonal to all the vectors from which it is composed. Swapping any two vectors negates Q. When N=2, one can rescue the otherwise failing vector notation by thinking of it as saying Q = B x . Comment: Notice that Q = B x C x D is defined for 4-vectors only. This is a completely different animal from the object Q = B x (C x D) which is defined for 3-vectors only. This latter object contains two ε factors, while the former only one. (d) Missing Man Formation We now make a small variation in the notation. Start with the above equation, Qa ≡ εabc...x BbCcDd.....Xx , then change a to α, back up all the Latin letters by one (but leave the last as "unknown" x), and assume that some subscript κ and factor Kκ are "missing". The result is, Qα ≡ εαac...x AaBbCc.....Xx // κ and Kκ are missing There are still N-1 factors, and one can still write this in vector notation Q = A x B x C x ... x X // K is missing and of course it is still true that QC = 0, etc. For N=2 the vector notation is rescued as in (c) above. (e) Apply this Notation to E Compare the above Qα to the section (a) definition of (Ek)α , (Ek)α ≡ det(R) (-1)k-1{ εαabc...x (e1)a(e2)b ...... (eN)x } // κ(k) and (ek)κ are missing; N≥ 2 Therefore, the definition of Eκ for N > 2 can be written in this vector notation, Ek ≡ det(R) (-1)k-1 e1 x e2 x ......x eN // ek missing; N > 2 The reciprocal base vector Ek is thus orthogonal to all the tangent base vectors from which it is constructed (remember ek is missing)! For example, for N=3 the three E vectors are given by E1 = det(R) (-1)1-1 e2 x e3 = det(R) e2 x e3 E2 = det(R) (-1)2-1 e1 x e3 = det(R) e3 x e1 E3 = det(R) (-1)3-1 e1 x e2 = det(R) e1 x e2 which agrees with the results quoted above. For N =2 ( E's label corresponds to the missing e's label ), E1 = det(R) (-1)1-1 e2 x = det(R) e2 x or (E1)k = det(R) εka(e2)a E2 = det(R) (-1)2-1 e1 x = - det(R) e1 x or (E2)k = -det(R) εka(e1)a One can combine these two lines into one as follows ( eg, k = 1, then 3-1 = 2, etc) Ek = det(R) (-1)k-1 e3-k x = det(R) e3-k x or (E1)k = det(R) (-1)k-1εka(e3-k)a The vector "trick" notation shows that E1e2 = 0 and E2e1 = 0, E1e2 = det(R) e2 x e2 = 0 E2e1 = -det(R) e1 x e1 = 0 and also E1e1 = det(R) εka(e2)a (e1)k = det(R)det[e1, e2] = det(R)det(S) = 1 E2e2 = -det(R) εka(e1)a (e2)k = -det(R)det[e2, e1] = det(R)det(S) = 1 It is shown next that these N=2 results are special cases of a general fact: Em en = δm,n . Section 5 (j) showed that em en = 'mn . The other two dot products are now considered. (f) Compute Em en One can now compute, for general N, Ek ek = (Ek)α(ek)α = { det(R) (-1)k-1εαabc...x (e1)a(e2)b ...... (eN)x } (ek)α k is missing Slide α to the right in the ε subscript field and put it into the hole of the missing subscript κ, picking up (-1)k-1. At the same time, move the (eκ)α to the left and position it in its proper place in the product of factors, Ek ek = (Ek)α(ek)α = { det(R) εabc..α..x (e1)a(e2)b ... (eκ)α ... (eN)x } = det(R) det [e1, e2, e3 .... eN ] = det(R) det(S) = 1 // since RS = 1 We already know that Ek is orthogonal to all the en which form the generalized cross product, therefore Em en = δm,n which is the "duality relation" discussed more generally in Section 6 (b) (g) Compute En Em Since the vectors { en } are linearly independent and thus form a basis in RN, Em can be expanded onto the en , Em = Σn An(m) en δm,k = Em ek = Σn An(m) en ek = Σn An(m) 'nk Multiplying both sides by g'ki and summing on k gives LHS = Σk g'ki δm,k = g'mi RHS = Σn An(m) (Σk 'nk g'ki) = Σn An(m) ('g')ni = Σn An(m)δn,i = Ai(m) Therefore Ai(m) = g'mi so, Em = Σn An(m) en = Σn g'mn en which is to say Ek is this linear combination of the ei (this is the definition used in Section 6 (a)) Ek = Σi g'ki ei = g'ki ei // implied sum on i // Std Notation: ek = Σi g'ki ei which may be compared with the previous result Ek ≡ det(R) (-1)k-1 e1 x e2 x ......x eN // ek missing; It seems rather impressive that these two dissimilar ways of writing E are equal. Finally, En Em = En (g'mi ei) = g'mi (En ei) = g'mi δn,i = g'mn = g'nm // recall g' symmetric (h) Summary of relationship between the tangent and reciprocal base vectors en em = 'nm En Em = g'nm en Em = δn,m En = Σi g'ni ei en = Σi 'ni Ei ' = g'-1 Although these results have just been derived in the Picture B context, they are also valid in the more general Picture A context, as shown in Section 6 in which the equation En = Σi g'ni ei is used as the definition of En . As a reminder, the cross product expression for En is only valid in Picture B. In Standard Notation, the summary above can be restated as en em = g'nm en em = g'nm en em = δnm en = Σi g'ni ei en = Σi g'ni ei g'ab = (g'ab) -1 (i) Another Cross Product Notation and another expression for E Go back to the general cross product of N-1 vectors each of dimension N, Q = B x C x D x ... x X // N-1 factors, N-2 crosses Replace B,C,D ... by vectors A(n), Q = A(1) x A(2) x A(3) x ... x A(N-1) // N-1 factors, N-2 crosses It is convenient to write this as Q = Πxi=1N-1 A(i) = Πxi A(i) where in the second form it is understood that i takes on all values i = 1 to N-1. The superscript x means that this is not a regular product, it is our generalized cross product. This Πx symbol also implies correct handling of the special case N=2 such that Q = Πxi=11 A(i) = A(1) x // ≠ A(1) as discussed in section (c) above. This same Πx notation can be applied to the "missing man formation" of (d) above. Suppose Q = A(1) x A(2) x A(3) x ... x A(N) // A(n) is missing One can write this as Q = Πxi=1..N,i≠n A(i) ≡ Πxi≠n A(i) And of course this idea can be applied to the expression for Ek Ek = det(R) (-1)k-1 e1 x e2 x ......x eN // ek missing; Ek = det(R) (-1)k-1 Πxi≠k ei Once again, for N=2 the Πx symbol implies that ( ek = "missing", e3-k = the one not missing) Πxi≠k ei = Πxi=1..2,i≠k ei = e3-k x Ek = det(R) (-1)k-1 e3-k x which is the "trick" notation of subsection (c) above for the N=2 case. Appendix B: The Geometry of Parallelepipeds in N dimensions Introduction This Appendix presents a simple method for constructing an N dimensional parallelepiped, which name we shorten to "N-piped". It is found that an N-piped has 2N vertices and N pairs of faces for a total of 2N faces, and the locus of points that make up each of these faces is stated. Each face of an N-piped is in fact an (N-1)-piped which has 2N-1 vertices and is planar in N dimensions (meaning it lies on an N-1 dimensional flat surface). The two faces which make up each face pair lie on parallel planes in RN. For example, for N=3 each face is a 2-piped having 23-1 = 4 vertices, and there are N=3 face pairs for a total of 6 faces, and each pair of faces is planar in 3 dimensions. For N=4 there are 4 pairs of faces for a total of 8 faces. Each face is a 3-piped having 24-1 = 8 vertices. For example, one would say that each face of a 4-cube is a 3-cube. It is not intuitively obvious that two faces each of which is a regular cube can in fact lie on surfaces which are planar and parallel in 4 dimensions, but we show how this works below. It is then shown that, if the N-piped is spanned by the N tangent base vectors en of Section 3, the normal vectors for the pairs of parallel faces are just the reciprocal base vectors En of Section 6. Section (d) focuses on the area and volume of N-pipeds in various dimensions, and simple expressions for the volume and vector areas of the faces of an N-piped are obtained. Rather than just state the results in N dimensions, we attempt an inductive approach to provide motivation for the N dimensional results. In this approach, cases N = 2,3.. are treated with nearly identical boilerplate templates to build up the inductive case. All major results of this Appendix are concisely stated in Summary section (e). Since this is a very long section (~15 p), a reader not interested in details would do well to simply read that summary and skip the rest of this Appendix. (a) Preliminary: Equation of a plane in N dimensions Consider an arbitrary plane drawn in N space which does not pass through the origin. There is some point on that plane which lies closer to the origin than all other points on the plane. Let p be a vector from the origin to that closest point, and let r represent a point lying on the plane, Since p is normal to the plane, and since r-p is a vector lying in the plane, it follows that p(r-p) = 0 => rp = p2 => r = p Therefore, one way to write the equation of a plane in N dimensions is r = p r = (x1, x2, .....xN) where is the unit vector normal to the plane which points "away from the origin", and where p > 0 is the distance of closest approach of the plane to the origin. In the limit p→0, the plane passes through the origin and the equation is then r = 0 where is either normal to the plane. (b) N-pipeds and their Faces in Various Dimensions The 1-piped Start with N = 1 where the piped is some arbitrary line segment e1 in direction 1 having length e1, with one end affixed to the origin of the real axis. N=1 rvolume1 = α1 e1 0 ≤ α1 ≤ 1 This piped has two vertices located at v1 = 0 and v2 = e1. These two vertices are also the "faces" of this 1-piped, so there are two faces (one pair of faces). These faces are 0 dimensional and therefore don't point in any direction (they are the endpoints of the segment). The 1-piped is a piece of a plane in 1 dimension (a line). One can think of the vertex at the origin as the "generator 0-piped" and the other vertex as the partner face of the generator, in the sense of the generator idea described below. The volume of this 1-piped is e1. The 2-piped Now add another dimension, going to N=2. Introduce a unit vector e2 in some arbitrary direction in R2 other than e1 so that e1 and e2 are linearly independent. Take the 1-piped described above (line segment) and translate it by e2 to create a new copy of the line segment. The original 1-piped we call the generator piped, and the copy is the partner of the generator piped which, it will be shown, lies on a plane (a 1-plane = line) which is parallel to the plane of the generator piped, but its plane does not pass through the origin. In N=2 dimensions, the generator 1-piped and its partner are now "faces" of a 2-dimensional object, a parallelogram = a 2-piped. Draw line segments from all the vertices of the generator piped to matching vertices of its partner piped (add 2 line segments) to make 2 additional "side" faces. One of these faces necessarily touches the origin, and the other face does not. Faces always occur in parallel pairs one of which touches the origin, and one of which does not, the latter we will call the "partner" face. For our 2-piped, each face is a 1-piped. There are now four faces, each is a line segment. The loci of the 2-piped's volume and of its four 1-piped faces are given by rvolume2 = α1e1 + α2 e2 0 ≤ α1,α2 ≤ 1 rface2 = α1e1 0 ≤ α1 ≤ 1 // the generator face rface2p = α1e1 + e2 0 ≤ α1 ≤ 1 // partner of the generator face rface1 = α2e2 0 ≤ α2 ≤ 1 // side face touching the origin rface1p = α2e2 + e1 0 ≤ α2 ≤ 1 // partner of the above side face The origin-touching faces are numbered using the index of the en vector that does not appear in the locus for the face. This seems strange but for N > 2 it will be clear why this is done. It is possible to construct vectors E1 and E2 as linear combinations of e1 and e2 such that the following is true (see Section 6 (b)) Ei ej = δi,j // Ek = Σi=12 g'ki ei , see Section 6 (a) If one interprets the en vectors as tangent base vectors for some transformation F, then the two vectors En are the corresponding reciprocal base vectors which are discussed in Section 6 and Appendix A. Consider now these dot products: E2 rface2 = E2 α1e1 = 0 => 2 rface2 = 0 E2 rface2p = E2 [ α1e1+ e2] = 1 => 2 rface2p = 1/E2 The first line says (section (a) above) that face 2 lies on a plane which passes through the origin and which has normal vector 2. The second line says that face 2p has the same normal and its plane is therefore parallel to face 1 but misses the origin by distance 1/|E1|. Similarly, E1 rface1 = E1 α2e2 = 0 => 1 rface1 = 0 E1 rface1p = E1 [ α2e2+ e1] = 1 => 1 rface1p = 1/E1 These two faces are also parallel, both having normal 1. The first touches the origin while the partner's plane misses the origin by distance 1/|E1| . The conclusions that En is normal to face n and that the pair of faces n and np are parallel do not depend on the specific upper endpoints of the ranges of α1 and α2 which happen to be given as 1 above. This seems pretty obvious since rescaling the edges of a parallelogram does not affect its normal vector. The 3-piped Now add another dimension, going to N=3. Introduce a unit vector 3 in some arbitrary direction in R3 so that (1,2,3) are linearly independent. Take the 2-piped described above (parallelogram) and translate it by distance e3 in the 3 direction to create a new copy of the 2-piped. The original 2-piped we call the generator piped, and the copy is the partner of the generator piped which, as will now be shown, lies on a plane which is parallel to that of the generator piped, but which does not pass through the origin. In N=3 dimensions, the generator 2-piped and its partner are now "faces" of a 3-dimensional object, a parallelepiped = a 3-piped. Draw line segments from all 22 vertices of the generator piped to the corresponding vertices of its partner piped (add 4 line segments), to get 4 additional side faces. Two of these faces necessarily touch the origin, and the other two do not. For the 3-piped, each face is a 2-piped. There are now 2*3 = 6 faces, each is a 2-piped. The loci of the 3-piped's volume and of its six 2-piped faces are given by rvolume3 = α1e1 + α2 e2 + α3 e3 0 ≤ α1,α2,α3 ≤ 1 rface3 = α1e1 + α2 e2 0 ≤ α1,α2 ≤ 1 // the generator face rface3p = α1e1 + α2 e2 + e3 0 ≤ α1,α2 ≤ 1 // partner face to the above rface2 = α1e1 + α3 e3 0 ≤ α1,α3 ≤ 1 // the generator face rface2p = α1e1 + α3 e3 + e2 0 ≤ α1,α3 ≤ 1 // partner face to the above rface1 = α2e2 + α3 e3 0 ≤ α2,α3 ≤ 1 // the generator face rface1p = α2e2 + α3 e3 + e1 0 ≤ α2,α3 ≤ 1 // partner face to the above Notice that the partner face locus is created from the non-partner face by adding "the other" base vector. For example, face 3 is "spanned" by base vectors e1 and e2 so e3 is added to get the partner. A partner is just a copy of the non-partner which is translated by a constant vector. The above results can be summarized in this concise manner: rvolume3 = Σnαnen 0 ≤ αn ≤ 1 rface(i) = Σn≠iαnen 0 ≤ αn ≤ 1 i = 1,2,3 rface(ip) = Σn≠iαne + ei 0 ≤ αn ≤ 1 i = 1,2,3 It is possible to construct vectors E1, E2, E3 as linear combinations of e1, e2, e3 such that the following is true (see Section 6 (b)) Ei ej = δij // Ek = Σi g'ki ei , see see Section 6 (a) If one interprets the en vectors as tangent base vectors, then the three vectors En are the corresponding reciprocal base vectors which are discussed in Section 6 and Appendix A. Consider now these dot products: E1 rface1 = E1 [ α2e2 + α3 e3] = 0 => 1 rface1 = 0 E1 rface1p = E1 [α2e2 + α3 e3 + e1] = 1 => 1 rface1p = 1/E1 The first line says that face 1 lies on a plane which passes through the origin and which has normal vector 1. The second line says that face 1p has the same normal and its plane is therefore parallel to face 1 but misses the origin by distance 1/|E1| A similar pair of equations obtains for each of the other face pairs. The N-piped Now add another dimension, going from N-1 to N. Introduce a unit vector N in some arbitrary direction in RN so that ( 1... N ) are linearly independent. Take the (N-1)-piped described above and translate it by distance eN in the N direction to create a new copy of the (N-1)-piped. The original (N-1)-piped we call the generator piped, and the copy is the partner of the generator piped which, it will be shown, lies on a plane which is parallel to that of the generator piped, but which does not pass through the origin. The generator (N-1)-piped and its partner are now "faces" of a N-dimensional object, an N-piped. Adding this partner piped doubles the total vertex count. Draw line segments from all 2N-1 vertices of the generator piped to the corresponding vertices of its partner piped to get 2N-2 additional side faces for a total now of 2N faces. There are N pairs of "faces" because there are N ways to omit a single ei from the list of vectors which span a face, so including the partner faces an N-piped has 2N faces in total. Half of these faces necessarily touch the origin, and the other half do not. Each face is an (N-1)-piped. It is convenient to refer to the partner face of a pair as "the far face" and the other one, which touches the origin, as "the near face". The loci of the N-piped's volume and of its 2N (N-1)-piped faces are given by: rvolumeN = Σnαnen 0 ≤ αn ≤ 1 rface(i) = Σn≠iαnen 0 ≤ αn ≤ 1 i = 1,2...N rface(ip) = Σn≠iαnen + ei 0 ≤ αn ≤ 1 i = 1,2...N Notice that the partner face locus is created from the non-partner face by adding "the other" base vector. For example, face i is "spanned" by base vectors en n ≠ i, so it is ei that one adds to get the partner. A partner is just a copy of the non-partner which is translated by a constant vector. It is possible to construct vectors E1...EN as linear combinations of e1.. eN such that the following is true (see Section 6 (b)) Ei ej = δij // Ek = Σi g'ki ei If one interprets the en vectors as tangent base vectors, then the N vectors En are the corresponding reciprocal base vectors which are discussed in Section 6 and Appendix A. Consider now these dot products: Ei rface(i) = Ei [Σn≠iαnen] = 0 => i rface(i) = 0 Ei rface(ip) = Ei [Σn≠iαne + ei] = 1 => i rface(ip) = 1/Ei The first line says that face i lies on a plane which passes through the origin and which has normal vector i. The second line says that face ip has the same normal and its plane is therefore parallel to face i but misses the origin by distance 1/|Ei| (c) The question of inward versus outward facing normal vectors. It has been shown above that, for an N-piped, the pair of faces i and ip has normal vector Ei. For one of these faces, Ei will be an outward directed normal, while for the other it will be an inward directed normal. One might like to know which is which. Here is one way to find out. First, construct these three vectors (piped)center = Σn(1/2)en // vector from origin to piped center (face i)center = Σn≠i(1/2)en // vector from origin to center of face i (face ip)center = Σn≠i(1/2) en + ei // vector from origin to center of face ip Construct vectors from piped center to face centers ( results here are fairly obvious) (face i)center - (piped)center = [Σn≠i(1/2)en] - Σn(1/2)en = - (1/2)ei (face ip)center - (piped)center = [Σn≠i(1/2)en+ ei ] - Σn(1/2)en = + (1/2)ei Then compute Ei {(face i)center - (piped)center } = Ei [- (1/2)ei ] = -(1/2) < 0 Ei {(face ip)center - (piped)center } = Ei [+ (1/2)ei ] = +(1/2) > 0 One may conclude that Ei is an outward pointing normal for face ip (far face). Therefore, - Ei is an outward pointing normal for face i, which recall is the face which touches the origin (near face). (d) The Face Area and Volume of N-pipeds in Various Dimensions We embark now on another long march to inductively arrive at results for the general N case. Tracing the first few cases N = 2,3,4 and then extrapolating to N = N probably gives more insight than a formal induction proof which is not attempted here. Each case below is treated with the same boilerplate template which first treats Face Area and then Volume. The 2-piped Face Area. The area of a 2-piped face (a line segment) is just the length of the edge which is the face, A1 = |e2| A2 = |e1| where here we maintain the plan of labeling an area by the index of the spanning vector which is omitted in making the area. The vector areas can be written, based on the work above, A1 = |e2| 1 // Ek = Σi g'ki ei , see Appendix A (g) A2 = |e1| 2 and these vectors are out-facing for faces 1p and 2p. We claim that both these results can be expressed in a single formula An = |det(S)| En . One can see that the direction is correct for n = 1,2, so it is just a question of verifying the magnitude. One must show that |det(S)| |E1| = |e2| and |det(S)| |E2| = |e1| or |Ek| = |det(R)| |e3-k| k=1,2 // RS = 1 Using the N=2 trick notation from Appendix A (c), Ek = det(R) (-1)k-1 e3-k x so that | Ek | = | det(R)| | e3-k x | = |det(R)| | e3-k| k = 1,2 since e3-k and are perpendicular. QED. We stress the formula An = |det(S)| En because it will turn out that this is valid for all N ≥ 2 . One can restate An = |det(S)| En using the cross product notation presented in Appendix A (i): An = |det(S)| En = |det(S)| det(R) (-1)n-1 Πxi≠n ei = σ (-1)n-1 Πxi≠n ei σ ≡ sign(det(S)) = sign(det(R)) Volume. The volume of a 2-piped is the base times the height of a parallelogram, familiarly given as by the cross product of the edges, volume(2) = | e1 x e2 | = | εab (e1)a(e2)b | = | det [ e1, e2 ] | = | det(S) | where S is the linearized transformation matrix for N=2, see Section 2. Of course strictly in N=2 the notation e1 x e2 has no meaning, so one has to imagine a dimension to give it meaning. The second form does have a meaning for N=2, and that meaning is |(e1)1 (e2)2 – (e1)2 (e2)1|. The 3-piped Face Area: The faces of a 3-piped are 2-pipeds. For N=2, the 2-piped volume was volume(2) = | εab (e1)a(e2)b | , where e1 and e2 were 2D vectors. For the 2-piped which is "face 3" of the 3-piped -- a "near" face which touches the origin of the 3D skewed en coordinate system -- vectors e1 and e2 are 3D vectors. The first 2 components of each of these 3D vectors are the same as the components of the 2D ei vectors, while the 3rd components are both 0. This is so because face 3 lies in a plane defined by this 3rd component being 0. The volume(2) formula expressed in terms of these new 3D vectors is therefore |εab3 (e1)a(e2)b|, where ε is now a 3D ε tensor. The conclusion is that A3 = |εab3 (e1)a(e2)b| and this then is the scalar area of both face 3 and its partner face 3p, the far face. Similar arguments would then support these other area expressions A1 = |ε1ab (e2)a(e3)b| A2 = |εa2b (e3)a(e1)b| Since indices on ε can be swapped for free due to the absolute value, the non-summed index can be put first in all three cases and one may then conclude that A1 = |e2 x e3| face 1 and face 1p A2 = |e3 x e1| face 2 and face 2p A3 = |e1 x e2| face 3 and face 3p In Appendix A (e) it was shown that E1 = det(R) e2 x e3 so that e2 x e3 lines up with E1. Regardless of the sign of det(R), we define the vector areas to point in the +n directions. Thus, A1 ≡ |e2 x e3| 1 face 1p out-facing A2 ≡ |e3 x e1| 2 face 2p out-facing A3 ≡ |e1 x e2| 3 face 3p out-facing These equations can be combined into the following single formula An = |e1 x ... x e3| n // en missing where the ei are reordered for free due to the absolute value signs. But Appendix A says En = det(R) (-1)n-1 e1 x ... x e3 // en missing so |En| = | det(R) | | e1 x ... x e3 | // en missing Thus, An = n |En| / |det(R)| = |det(S)| En = |det(S)| det(R) (-1)n-1e1 x ... x e3 // en missing = σ (-1)n-1e1 x ... x e3 // en missing where σ ≡ sign[det(S)] = sign[det(R)] To summarize, for N=3 one has An = |det(S)| En = σ (-1)n-1e1 x ... x e3 // en missing = σ (-1)n-1 Πxi≠n ei where the last line uses the shorthand notation of Appendix A (i). These expressions have the same form as those of the 2-piped. Volume. The volume of a 3-piped (with each of the above An in turn treated as the base) is base times height, so volume(3) = | A1 e1 | = | A2 e2 | = | A3 e3 | or volume(3) = | e2 x e3 e1 | = | e3 x e1 e2 | = | e1 x e2 e3 | Here is a drawing showing the last case (σ = +1), where "base" is A3 = | e1 x e2 | and "height" is e3 cosθ , Using ε notation one can write e3 e1 x e2 = (e3)i εijk(e1)j(e2)k = εijk (e1)j(e2)k (e3)i = εjki (e1)j(e2)k(e3)i = det [ e1, e2, e3] = det(S) so that volume(3) = | e3 e1 x e2 | = | εabc (e1)a(e2)b(e3)c | = | det [ e1, e2, e3] | = | det(S) | These expressions have the same form as those of the 2-piped. The 4-piped Face Area: The faces of a 4-piped are 3-pipeds. For N=3, the 3-piped volume was volume(3) = | εabc (e1)a(e2)b(e3)c | where e1,e2,e3 were 3D vectors. For the 3-piped which is "face 4" of the 4-piped -- a "near" face which touches the origin of the 4D skewed en coordinate system -- vectors e1,e2,e3 are 4D vectors. The first 3 components of each of these 4D vectors are the same as the components of the 3D ei vectors, while the 4th components are all 0. This is so because face 4 lies in a plane defined by this 4th component being 0. The volume(3) formula expressed in terms of these new 4D vectors is therefore | εabc4 (e1)a(e2)b(e3)c |, where ε is now a 4D ε tensor. The conclusion is that A4 = | εabc4 (e1)a(e2)b(e3)c | and this then is the scalar area of both face 4 and its partner face 4p, the far face. Similar arguments would then support these other area expressions A1 = | ε1abc (e2)a(e3)b(e4)c | A2 = | εa2bc (e3)a(e4)b(e1)c | A3 = | εab3c (e4)a(e1)b(e2)c | Since indices on ε can be swapped for free due to the absolute value, the non-summed index can be put first in all three cases and one may then conclude that A1 = |e2 x e3 x e4| face 1 and face 1p A2 = |e3 x e4 x e1| face 2 and face 2p A3 = |e4 x e1 x e2| face 3 and face 3p A4 = |e1 x e2 x e3| face 4 and face 4p where, as discussed in Appendix A (c), Q = A x B x C is defined by Qk = εkabc AaBbCc . In Appendix A (e) it was shown that E1 = det(R) e2 x e3 x e4 so that e2 x e3 x e4 lines up with E1. Regardless of the sign of det(R), we define the vector areas to point in the +n directions. Thus An = |e1 x ... x e3| n // en missing n = 1,2,3,4 where the ei are reordered for free due to the absolute value signs. But Appendix A says En = det(R) (-1)n-1 e1 x ... x e4 // en missing so |En| = | det(R) | | e1 x ... x e4 | // en missing Thus, An = |En| n / |det(R)| = |det(S)| En = |det(S)| det(R) (-1)n-1e1 x ... x e4 // en missing = σ (-1)n-1e1 x ... x e4 // en missing σ ≡ sign[det(S)] = sign[det(R)] To summarize, for N=4 one has An = |det(S)| En = σ (-1)n-1e1 x ... x e4 // en missing = σ (-1)n-1 Πxi≠n ei σ ≡ sign[det(S)] = sign[det(R)] These expressions have the same form as those of the 2-piped and the 3-piped. Volume. The volume of a 4-piped (with each of the above An in turn treated as the base) is base times height, so volume(4) = | A1 e1 | = | A2 e2 | = | A3 e3 | = | A4 e4 | or volume(4) = | e2 x e3 x e4 e1 | = | e3 x e4 x e1 e2 | = | e4 x e4 x e1 e3 | = | e1 x e4 x e2 e4 | Using ε notation one can write the first case as e2 x e3 x e4 e1 = (e1)a εabcd(e2)b(e3)c(e4)d = εabcd(e1)a(e2)b(e3)c(e4)d = det [ e1, e2, e3, e4] = det(S) so that volume(4) = | det(S) | = | det [ e1, e2, e3, e4] | = | εabcd(e1)a(e2)b(e3)c(e4)d | These expressions have the same form as those of the 2-piped and the 3-piped. The N-piped Face Area: The faces of a N-piped are (N-1)-pipeds. If there had been an (N-1)-piped section prior to this one, the volume formula there would have been volume(N-1) = | εabc...x (e1)a(e2)b... (eN-1)x | where e1,e2,...eN-1 were (N-1)D vectors. For the N-piped which is "face N" of the N-piped -- a "near" face which touches the origin of the ND skewed en coordinate system -- vectors e1,e2,...eN-1 are ND vectors. The first N-1 components of each of these ND vectors are the same as the components of the (N-1)D ei vectors, while the Nth components are all 0. This is so because face N lies in a plane defined by this Nth component being 0. The volume(N-1) formula expressed in terms of these new ND vectors is therefore | εabc...xN (e1)a(e2)b... (eN-1)x |, where ε is now an ND ε tensor. The conclusion is that AN = | εabc...xN (e1)a(e2)b... (eN-1)x | and this then is the scalar area of both face N and its partner face Np, the far face. Similar arguments would then support similar expressions for the other faces, for example, A1 = | ε1abc...x (e2)a(e3)b... (eN-1)w (eN)x | A2 = | εa2bc...x (e3)a(e4)b... (eN)w (e1)x | Since indices on ε can be swapped for free due to the absolute value, the non-summed index can be put first in all three cases and one may then conclude that A1 = |e2 x e3 x e4 x e5... x eN| face 1 and face 1p A2 = |e3 x e4 x e5 x e6... x e1| face 2 and face 2p A3 = |e4 x e5 x e6 x e7... x e2| face 3 and face 3p ... AN = |e5 x e6 x e7 x e8.. x e3| face N and face Np where, as discussed in Appendix A (c), Q = A x B x C.... X is defined by Qk = εkabc...x AaBbCc .....Xx In Appendix A (e) it was shown that E1 = det(R) e2 x e3 ... eN so that e2 x e3 ... eN lines up with E1. Regardless of the sign of det(R), we define the vector areas to point in the +n directions. Thus An = |e1 x ... x eN| n // en missing n = 1,2,3...N where the ei are reordered for free due to the absolute value signs. But Appendix A says En = det(R) (-1)n-1 e1 x ... x eN // en missing so |En| = | det(R) | | e1 x ... x eN | // en missing Thus, An = |En| n / |det(R)| = |det(S)| En = |det(S)| det(R) (-1)n-1e1 x ... x eN // en missing = σ (-1)n-1e1 x ... x eN // en missing σ ≡ sign[det(S)] = sign[det(R)] To summarize, for N=N one has An = |det(S)| En = σ (-1)n-1e1 x ... x eN // en missing = σ (-1)n-1 Πxi≠n ei σ ≡ sign[det(S)] = sign[det(R)] These expressions have the same form as those of the 2-piped, the 3-piped and the 4-piped. Volume. The volume of a N-piped (with each of the above An in turn treated as the base) is base times height, so volume(N) = | A1 e1 | = | A2 e2 | = ... = | AN eN | or volume(N) = | e2 x e3 x e4....eN e1 | = ... Using ε notation one can write the first case as e2 x e3 x e4...eN e1 = (e1)a εabc...x(e2)b(e3)c.....(eN)x = εabc...x(e1)a(e2)b(e3)c....(eN)x = det [ e1, e2, e3, ....eN] = det(S) where x is the Nth letter of the alphabet, so that volume(N) = | det(S) | = | det [ e1, e2, e3, ....eN] | = | εabc...x(e1)a(e2)b(e3)c....(eN)x | These expressions have the same form as those of the 2-piped, the 3-piped and the 4-piped. This result is also consistent with the volume(N-1) expression stated above. (e) Summary of Main Results of this Appendix 1. One way to write the equation of a plane in N dimensions is r = p r = (x1, x2, .....xn) where is the unit vector normal to the plane which points "away from the origin", and where p > 0 is the distance of closest approach of the plane to the origin. In the limit p→0, the plane passes through the origin and the equation is then r = 0 where is either normal to the plane. 2. An N-piped has 2N vertices as demonstrated by the inductive construction method presented above. 3. The locus of points making up the (closed) interior of an N-piped spanned by e1...eN is given by rvolumeN = Σn=1Nαnen 0 ≤ αn ≤ 1 The tails of all the vectors e1...eN meet at the origin of RN space. 4. There are N pairs of faces on an N-piped, and each face is an (N-1)-piped having 2N-1 vertices. The total face count is 2N. Each face is spanned by a subset of N-1 of the base vectors en, so each face is "missing" one of the en and the face is labeled using the index of this missing base vector. The loci of points making up the faces of an N-piped are given by rface(i) = Σn≠iαnen 0 ≤ αn ≤ 1 i = 1,2...N rface(ip) = Σn≠iαnen + ei 0 ≤ αn ≤ 1 i = 1,2...N where "face i" has a corner touching the origin of the N-piped (near face), while its parallel partner face "ip" does not touch the origin (far face). 5. If the N-piped spanning vectors en are the tangent base vectors associated with some transformation F, then Ei ej = δi,j where Ei are the reciprocal base vectors. In this case, the equations of the faces of the N-piped can be written in the form shown in item 1 above, i r(face i) = 0 i r(face ip) = 1/Ei i = 1,2...N so that both faces of a pair i are planar (in N dimensional space) and they have the same normal vector i so the faces of a pair lie on parallel planes. 6. The vector Ei is an outward-facing normal for face ip, while -Ei is an outward-facing normal vector for face i (which touches the origin). 7. The out-facing vector area of face ip of an N-piped can be expressed as Ai = |det(S)| Ei Ai = σ (-1)i-1 Πxj≠i ej σ ≡ sign[det(S)] = sign[det(R)] Ai = σ (-1)i-1 e1 x e2 ... x eN // ei missing where ei is the vector missing from the face's spanning set. The outfacing area for face i is - Ai. The last two lines are shorthands for the following, as discussed in Appendix A (i), (Ai )α = σ (-1)i-1 εαabc..x (e1)a (e2)b ... (eN)x // where (ei)ί and ί are missing For N=2 the last two expressions for Ai are interpreted as shown in Appendix A (c) Ai = σ (-1)i-1 e3-i x i = 1,2 8. The volume of an N-piped spanned by e1...eN is given by volume(N) = | det(S) | = | det [ e1, e2, e3, ....eN] | = | εabc...x(e1)a(e2)b(e3)c....(eN)x | where one can regard the tangent base vectors as the columns of the linearized transformation matrix S. Appendix C: Elliptical Polar Coordinates ( N=2, non-orthogonal) This Appendix is written in the developmental notation of Sections 1-6. (a) Elliptical polar coordinates The 2D "elliptic" coordinate system has coordinate lines which are orthogonal ellipses and hyperbolas. When rotated about its two symmetry axes, this system generates 3D prolate or oblate spheroidal coordinates. This is not the 2D coordinate system described in this Appendix. For "elliptical polar" coordinates, the coordinate lines are taken instead as the ellipses from elliptic coordinates, and the rays from polar coordinates. This non-orthogonal system is perhaps not very useful, but provides a good "sandbox" in which to study general aspects of coordinate systems. The transformation x' = F(x) is given by x'-space x-space (Cartesian) ρ2 = x2/a2 + y2/b2 x2+ y2 = r2 still x'1 = θ x1= x tanθ = y/x x'2 = ρ x2 = y Writing the first equation above as 1 = x2/(ρa)2 + y2/(ρb)2 it should be clear that ρ serves to label an ellipse of semi-major axis ρa, and semi-minor axis ρb, while θ labels the ray at angle θ, as in polar coordinates. The inverse transform x = F-1(x') is given by x = aρcosθ x/a = ρcosθ => x2/a2 + y2/b2 = ρ2 y = bρsinθ y/b = ρsinθ => tanθ = y/x The matrix S is given by S11 = (∂x/∂θ) = -aρsinθ S12 = (∂x/∂ρ) = acosθ Sik ≡ ( ∂xi/∂x'k) S21 = (∂y/∂θ) = bρcosθ S22 = (∂y/∂ρ) = bsinθ S = => det(S) = -abρ and R = S-1 = The tangent base vectors en can be read off as the columns of S e1 = ρ(-asinθ,bcosθ) = eθ |eθ| = ρ ≡ hθ eθ = |eθ| θ e2 = (acosθ,bsinθ) = eρ |eρ| = ≡ hρ eρ = |eρ| ρ The covariant metric tensor is, ' = STS = = which is clearly non-diagonal (but symmetric) as expected. When a = b = 1 it reduces to the polar coordinates system metric tensor where then ρ = r. The coordinate system is non-orthogonal because e1e2 ≠ 0, or equivalently, because ' is non-diagonal. (b) Forward coordinate lines Here is a Maple plot of some x-space forward coordinate lines (parameters a = 2 and b = 1) The coordinate lines in x-space are plotted using these equations, y = b // ellipses ρi= 1,2...10 10 ellipses y = x tanθi // rays θi = 2π (i/20) , i = 1,2...20 20 rays which are obtained from the forward transformation equations ρ2 = x2/a2 + y2/b2 tanθ = y/x (c) Inverse coordinate lines Here is a Maple plot of some x'-space inverse coordinate lines (parameters a = 2 and b = 1) The coordinate lines in x'-space are plotted using these equations ρ = xi/(acosθ) xi = -10 to +10 21 blue curves( one is a boxy U ) ρ = yi/(bsinθ) yi = -10 to +10 21 red curves ( one is a boxy U) which are obtained from the inverse transformation equations x = aρ cosθ y = bρ sinθ The secθ and cscθ curve families appear to "change shape", but that is just what happens when functions are scaled up vertically but not horizontally. If one plots one sine hump at different vertical scalings, the humps have different shapes. (d) Drawing a contravariant vector V in x-space: the meaning of V'n . A contravariant vector field V(x) can be expanded in these two ways (Section 6 (f)) V = V1(x) + V2(x) = Vx(x) + Vy(x) // un = for Cartesian V = V'1(x') e1 + V'2(x') e2 = V'θ(x') eθ + V'ρ(x') eρ V'(x') = R(x) V(x) where the V'n are the components of V transformed into x'-space where V becomes V'. The prime is not necessary on V'ρ but we maintain it as a reminder that it is an x'-space component. R(x) is the matrix of Section 2 and en are the tangent base vectors of Section 3. The fields V'n(x') are "the components of vector field V' in x'-space", since V' = RV , or V'i = RijVj. Moreover, these V'n(x') are "expressed in terms of curvilinear coordinates" (Section 2(i)). If one is asked to "express a vector V in curvilinear coordinates", one is usually being asked to write V as the second expansion above. The vectors en and V exist in x-space, and in the second expansion it just happens that the coefficients V'n(x') are the components of V', which is a vector in x'-space, and are functions of the curvilinear coordinates. Here is a graphical representation of this vector V in x-space: As advertised, the tangent base vectors are not at right angles. The V parallelogram accurately illustrates the equation V = V'θ eθ + V'ρ eρ. Since x-space is Cartesian, there is no distinction between Cartesian length (graphical length) and covariant length for vectors in x-space. Graphically, one could find the values for V'θ and V'ρ as follows: (1) for the point (x,y), compute the vectors eθ and eρ and compute their lengths |eθ| = h'θ and |eρ| = h'ρ ; (2) draw the parallelogram shown aligned with these vectors for some given V and find the edge lengths. The edges of the parallelogram are V'θ h'θ and V'ρ h'ρ so then the values of V'θ and V'ρ can be found. The alternative method is to compute R and use V'i = RijVj. (e) Drawing a contravariant vector V' in x'-space: two "Views" As with previous examples, the above picture is drawn to the right of an x'-space picture as follows: In Section 3 the axis-aligned basis vectors e'n were introduced as e'n , n = 1,2...N // (e'n)i = δn,i e'1 = (1,0,0...) etc and en was shown to be a contravariant vector, e'n = R(x) en. Applying matrix R(x) to the equation V = V'θ eθ + V'ρ eρ one gets V' = V'θ e'θ + V'ρ e'ρ which appears first in the list of expansions of V' in Section 6 (f). There is no ambiguity concerning this last equation. Ambiguity can arise, however, when one tries to represent this equation graphically in x'-space. There are two very different "views" one can take of a drawing in x'-space. In the first view, we take x'-space to be "flat" (Cartesian) so that g' = 1. In the second view, we take x'-space to be "curved" with g' ≠ 1. These views are really two different x'-spaces since the metric tensors are different. In the Cartesian View of x'-space the length (norm) of a vector is given by |A|2 = δijAiAj = ΣAi2, so one has, since (e'n)i = δn,i, ' = 1 |e'n| = 1 e'n = 'n = ' = the usual axis-aligned unit vectors in x'-space V' = V'θ + V'ρ = '1 = '2 The left-side graph shown above is, in this view, a "normal Cartesian graph" and the vectors add up properly, for example Pythagoras tells us that |V'|2 = V'θ2 + V'ρ2 and = 1 = 1 = 0 This Cartesian View, which is x'-space with ' = 1, is appropriate in applications in which it is not required or desired that norms and dot products be tensorial scalars, as discussed at the end of Section 5 (a). For example, when ' is set to 1, one has |V'| ≠ |V| in the above picture. Another use of this view involves integration as will be seen below. In the Curvilinear View of x'-space, one assumes that ' takes a value which enforces the scalarity of norms and dot products between x-space and x'-space, which is to say, one takes ' = ST S where is the x-space metric tensor for x-space. Normally =1 (Cartesian x-space), so '= STS. In this Curvilinear View, then, the length |A'| of a contravariant vector A' is determined by |A'|2 = 'ijA'iA'j where ' = STS ≠ 1 |e'n| = |en| = h'n ≡ n = 1,2 for θ,ρ V' = V'θ e'θ + V'ρ e'ρ = V'θ h'θ 'θ + V'ρ h'ρ 'ρ 'n ≡ e'n |en| = e'n h'n |V'|2 = |V|2 = Vx2 + Vy2 ≠ (V'θ h'θ)2 + (V'ρ h'ρ)2 // unless x'i are orthogonal coordinates This last inequality says that in the Curvilinear View the Pythagorean Theorem is invalid. In fact |V'|2 = 'ijV'iV'j = 'θθ V'θ2 + 'ρρ V'ρ2 + 2 'θρV'θ V'ρ = (V'θ h'θ)2 + (V'ρ h'ρ)2 + 2 'θρV'θ V'ρ In writing |e'n| = |en| and |V'|2 = |V|2 above, we use the rule shown in Section 5 (i) which says |A'|2 = |A|2 for any contravariant vector A (|A|2 is a scalar ). Moreover, 'θ 'θ = e'θ e'θ / (h'θ2) = eθ eθ / (h'θ2) = 'θθ / (h'θ2) = 1 'ρ 'ρ = e'ρ e'ρ / (h'ρ2) = eρ eρ / (h'ρ2) = 'ρρ / (h'ρ2) = 1 'θ 'ρ = e'θ e'ρ / (h'θh'ρ) = eθ eρ / (h'θh'ρ) = 'θρ / (h'θh'ρ) ≠ 0 <= !! so that the 'n are unit vectors having unit covariant length, but 'θ 'ρ ≠ 0 despite the fact that these vectors are drawn at right angles in the x'-space graph above, en = h'n 'n. One might imagine trying to slant the lines of the x'-space graph to cause all intersection points to have angles which match the metric tensor, which is to say, at each intersection point one would need an angle ψ where 'θ 'ρ = cosψ. But in general 'n 'm = 'nm / (h'nh'm) has a different value at every point, so such a graph would be quite complex. The upshot is that for a non-orthogonal system, the axes in x'-space are still drawn at right angles and the purpose of the graph is mainly to "locate" all the points x' which correspond to points x in x-space according to x' = F(x). The graph does successfully represent the idea that V' = V'ρ e'ρ + V'θ e'θ, but one must give up on Euclidean geometry for this vector sum triangle. It might be imagined that the x'-space graph is the projection onto the plane of paper of some vectors drawn on a curved surface emerging from the plane of paper, and that is then why Pythagoras is wrong. In the case of an orthogonal coordinate system (diagonal '), the 90 degree angles between the axes in x'-space are accurate representations of the fact that 'n 'm = 0 when n ≠m. And since scalars are preserved, one has in the Curvilinear View, |V'|2 = Σn (h'nV'n)2 = Σn V'n2 = |V|2 = Σn Vn2 // orthogonal only where V'n ≡ h'nV'n and V' =Σn V'n 'n and V =Σn V'n n One can then still apply regular Euclidean geometry to the vector addition N-piped in x'-space in the sense that |V'|2 = Σn (h'nV'n)2. (f) Drawing the specific contravariant vector dx in x-space and x'-space Since dx is the primordial contravariant vector, everything stated in the last two sections applies with V → dx and V'θ → dx'θ = dθ, V'ρ → dx'ρ = dρ, where we finally drop the primes on dθ and dρ. The expansions of dx and dx' are, dx = dθ eθ + dρ eρ // in x-space dx' = dθ e'θ + dρ e'ρ // in x'-space For V = dx the picture above becomes It must be understood that now the vector arrows like dx are highly magnified and in reality are very small compared to, say, the curvature of the ellipse. From above, e'θ e'θ = 'θθ 'θ 'θ = 1 e'ρ e'ρ = 'ρρ 'ρ 'ρ = 1 e'ρ e'θ = 'ρθ 'θ 'ρ = 'θρ / (h'θh'ρ) and once again the "right angle" in the x'-space picture is deceptive. (g) Study of how dx transforms in the mapping between x-space and x'-space Consider this drawing which shows a representative set of vectors dx in x-space (the bars), along with the forward mappings (dx' = F(dx) or dx' = Rdx ) of the corresponding vectors dx' in x'-space. The vectors on the right all point up, those on the left point generally to the northeast. x'-space x-space Now select the red dx bar on the right and operationally apply the previous picture. First determine the tangent base vectors eθ and eρ at the location of the red bar. Then setting dx = dθ eθ + dρ eρ, consider the value of the two numbers dθ and dρ for this red bar. Graphically, knowing which way eθ and eρ point at the bottom of the red dx, one expects dθ > 0 and dρ > 0. The red dx' bar on the left has these Cartesian values dθ and dρ, and has a Cartesian-view length of |dx|2 = (dθ)2+(dρ)2. One can see from the picture that these Cartesian lengths vary for the 10 bars shown, though the lengths are all the same in x-space. The Curvilinear-view lengths of the x'-space bars are all the same, and are equal to the Cartesian length of those bars in x-space since dx'dx' = dxdx. Consider now some bar mapping in the other direction: Now the dx bars on the right all have different lengths. Those on the left have the same Cartesian length, which is what the drawing shows, but each one's Curvilinear-view length matches that of its corresponding bar on the right. The ratio of the length of a bar on the right to the Cartesian length of the corresponding bar on the left is the scale factor hθ which recall is a function of location in space: bar on right = dx(1) = e1 dx'1 = 1 h'1 dx'1 = eθ dθ = θ hθ dθ graph length = hθ dθ bar on left (Cartesian view) = dx'(1) = e'1 dx'1 = '1 dx'1 = θ dθ graph length = dθ => right bar length / left bar length = hθ = ρ (increases with ρ) If a=b, then hθ = ρ and the bar length on the right is then ρdθ as is obvious in polar coordinates. (h) Derivation of the Jacobian Integration Rule Consider now an integral ∫dθdρ f(θ,ρ). The tiny rectangles of area dθdρ, like the specific gray and orange ones highlighted on the left above, are regarded for the purposes of integration as being in the Cartesian view of x'-space. One then writes [ dA' is called dV' in Section 8 ] dA' ≡ dρdθ = the area of a differential patch in Cartesian-view x'-space This is the graphical area one sees in the picture. There is no need to define or consider any Curvilinear-view area in x'-space because the Cartesian-view area is being used. In the limiting process which defines the integration, each dθdρ patch on the left has the same area dρdθ. The interior of each patch on the left maps into some parallelogram patch on the right. One is not surprised to see that the patch areas on the right are different, though they map into patches on the left of the same Cartesian-view area. As shown in Section 8 (e), the ratio of the two patch areas is the absolute value of the Jacobian |J(x')|, (area of skewed patch on the right at location x) = |J(x')| dA' = |J(x')| dρdθ This is not what we mean by "the Jacobian Integration Rule" in the section title. That is coming below and it is going to involve the quantity dxdy. The mapping shown above between patches is an N=2 example of the general N-dimensional discussion in Section 8 (a) which describes an orthogonal differential N-piped in (Cartesian-view) x'-space mapping into a non-orthogonal differential N-piped in x-space. Now back to the integration issue. There are two ways an integration can be done in Cartesian x-space: integral of f(x) = lim Σi dA1(xi) f(xi) dA1(xi) = patches shown on the right above integral of f(x) = lim Σi dA2(xi) f(xi) dA2(xi) = dxdy In the first integral, every patch dA1(xi) on the right has a different shape and a different area as the integral is computed in the usual limiting-sum manner. The gray and orange patches on the right are two of these many patches. Despite their non-uniform shape and area, this rag-tag band of patches certainly "covers" the area being integrated over, and does so perfectly in the calculus limit. The area of one of these rag-tag patches is |J(x')|dA' = |J(x')|dθdρ and the areas are different because the Jacobian is a function of x = x(x'). In the second integral, every patch dA2(xi) has the same area dxdy, so really dA2(xi) does not depend on xi in this form of the integration. One such dxdy patch is shown in green above. The coverage of the dA2 patches is of course also "perfect coverage" in the calculus limit. Since both integrals cover the same area perfectly, they both give the same result in the limiting process that defines the integral. This point is sometimes misunderstood. One is not just "replacing" a parallelogram patch such as the orange one on the right with some dxdy patch that approximates it in area, like the green patch. The statement is about an integration. Thus one has lim Σi dA1(xi) f(xi) = lim Σi dA2(xi) f(xi) or ∫[|J(x')| dθdρ] f(x(x')) = ∫[dxdy] f(x) where on the left f(x) = f(x(x')) where x = F-1(x') ≡ x(x'). In the sense of distribution theory (Stakgold Chapters 1 and 5), one can then make this symbolic statement |J(θ,ρ)| dρdθ = dxdy where the meaning of this symbolic equality is the integral statement above, ∫D dxdy f(x) = ∫D' dθdρ |J(x')| f(x(x')) , valid for any integrable f(x) and any integration region D (region D' corresponds to D in x'-space.) Either of these last two equations constitute the "Jacobian Integration Rule" of the section title. The integral on the left is well defined in 2D calculus, so the expression on the right shows how to "evaluate the integral on the left in curvilinear coordinates". At this point one may introduce a new but obvious symbol dA ≡ dxdy so the above equality of integrals can be written ∫dA f(x) = ∫ dA' |J(x')| f(x(x')) |J(x')| dA' = dA In N dimensions, dA and dA' are differential "volumes", and the general Jacobian Integration rule takes the form ∫dV f(x) = ∫ dV' |J(x')| f(x(x')) |J(x')| dV' = dV dV' ≡ dx'1dx'2....dx'N = the volume of an orthogonal differential N-piped in the Cartesian-view x'-space dV = dx1dx2....dxN = the volume of an orthogonal differential N-piped in x-space. Notice that these are not the two N-pipeds which "map into each other" as noted above. The N-piped dV has nothing to do that that mapping which involved a non-orthogonal N-piped in x-space. To finish off our sample N=2 case, recall from earlier that for our polar elliptical coordinate system |J'(x')| = | det(S)| = abρ and therefore ∫dxdy f(x,y) = ∫ dθdρ |J(x')| f(aρ cosθ, bρ sinθ) = ab ∫dθdρ ρ f(aρ cosθ, bρ sinθ) In the limit of regular polar coordinates, one then has a = b = 1 and ρ = r so ∫dxdy f(x,y) = ∫rdrdθ f(rcosθ, rsinθ) which is the familiar result. Appendix D. Tensor Densities and the ε tensor Picture A is used in this Appendix along with Standard Notation. (a) Definition of a tensor density First, recall from the Section 5 (k) discussion of the Jacobian J, J ≡ det(Sij) = σ / = σ(sg'/sg)1/2 = σ(g'/g)1/2 => (g'/g)1/2 = σJ = |J| > 0 s = sign[det(gij)] = sign(g) = sign(g') g = det(gij) sg = |g| > 0 Sij ≡ (∂xi/∂x'j) σ = sign[det(Sij)] = sign(J) g' = det(g'ij) sg' = |g'| > 0 For proper Lorentz transformations of special relativity, det(S) = 1 so σ = +1. For curvilinear coordinates, one normally selects an ordering of the xi so that σ = +1, such as r,θ,φ in spherical coordinates. Nevertheless, we allow for the possibility of J < 0. Second, recall our generic sample tensor transformation from Section 7 (j) T' abcde = Raa' Rbb' Rcc' Sd'd Se'e Ta'b'c'd'e' which can be rewritten in a more standard way using the theorem of Section 7 (q) that Sμν = Rνμ, T' abcde = Raa' Rbb' Rcc' Rdd' Ree' Ta'b'c'd'e' T is a mixed rank-5 tensor, meaning it transforms as shown above with respect to the underlying transformation F. T is an "ordinary" standard-issue tensorial tensor. Now suppose instead that the object T were to transform like this, T' abcde = J-W Raa' Rbb' Rcc' Rdd' Ree' Ta'b'c'd'e' where the extra factor J-W has been introduced. If T transforms this way, it is called a tensor density of weight W. Thus, an ordinary tensor is a tensor density of weight 0. The convention for the sign of W used here is that of Weinberg p 99 Eq. (4.4.4), which equation has the following factor on the right side of a sample tensor density transform equation |∂x'/∂x|W ≡ [det(∂x'/∂x)]+W = [ det(∂x'i/∂xk) ]+W = J-W Some authors use -W as the "weight" instead of +W, but we shall stick with Weinberg's convention. An immediate example of a tensor density is provided by (g'/g)1/2 = |J| rewritten as g' = J2 g =J-(-2) g => weight(g) = -2 g' is a scalar density of weight - 2 This is the scalar density mentioned in Section 5 (k) of weight -2. Notice that from g one can construct other scalar densities of other weights, for example g'-1 = J-(2) g-1 => weight(g-1) = +2 g'-1 is a scalar density of weight +2 (b) A few facts about tensor densities 1. It is pretty obvious that a sum of two index-similar tensor densities of weight W has weight W. 2. Contracting indices within a tensor does not alter its weight W. If indices a and d are contracted in the example above, one gets T' abcae = J-W Raa' Rbb' Rcc' Rad' Ree' Ta'b'c'd'e' = J-W (Raa'Rad') Rbb' Rcc' Ree' Ta'b'c'd'e' = J-W δa'd' Rbb' Rcc' Ree' Ta'b'c'd'e' = J-W Rbb' Rcc' Ree' Ta'b'c'a'e' The factor J-W just sits there, impervious to contraction activities. 3. Going the other direction, when a larger tensor density is formed from two smaller ones, called a direct product or outer product, the weights get added. For example, A'a = J-W1 Raa'Aa' B'cd = J-W2 Rcc'Rdd' Bc'd' => (A'a B'cd) = J-(W1+W2) Raa' Rcc'Rdd' (Aa' Bc'd') 4. Although sometimes authors take a differing stance for certain tensors, we shall assume that indices are raised and lowered on a tensor density in exactly the same way they are raised and lowered on an ordinary tensor of the same index structure. This means the gab raises an index and gab lowers an index. 5. Raising or lowering an index does not change the weight of a tensor density. Again, using our generic example above, T'abcde = J-W Raa' Rbb' Rcc' Rdd' Ree' Ta'b'c'd'e' T'abcde = g'ex T'abcdx // raise last index in x'-space Ta'b'c'd'e' = ge'e" Ta'b'c'd'e" // lower last index in x-space Therefore T'abcde = g'ex [ J-W Raa' Rbb' Rcc' Rdd' Rxe' Ta'b'c'd'e'] = g'ex [ J-W Raa' Rbb' Rcc' Rdd' Rxe' (ge'e" Ta'b'c'd'e")] = JW Raa' Rbb' Rcc' Rdd' (g'ex Rxe' ge'e") Ta'b'c'd'e" = J-W Raa' Rbb' Rcc' Rdd' (Ree") Ta'b'c'd'e" // Section 7 (o) and again J-W passively watches all the action fly by. The weight of our generic tensor density with its last index raised is still W. 6. The covariant dot product of vector densities A and B of weights W and w is a scalar density of weight W + w and therefore A'B' = J-(W+w) AB . Proof: First form the rank-2 tensor density AiBj which by item 3 has weight W+w. Lower the second index and the mixed rank-2 tensor AiBj by item 5 still has weight W+w. Then contract to get AB = AiBi and by item 2, the weight is still W+w. Corollary: The magnitude of a vector density A of weight W is a scalar density of weight W, and therefore |A|' = J-W |A| Proof: |A|2 = A A which has weight 2W meaning |A'|2 = J-2W |A|2. Therefore |A'| = J-W |A| . 7. As J→1, tensor densities become true tensors. One could imagine some limiting/morphing process on an underlying transformation F such that the linearized transformation matrix R approaches a rotation matrix at all points in space (RRT= 1 and detR = 1) and then J = detS → 1. In this case J-W → 1-W = 1 and therefore any tensor density, regardless of its weight W, becomes an ordinary tensor. Perhaps we should restrict this comment to underlying transformations F having detS > 0 since passing through detS = 0 is problematical. An example: the cross product considered in section (g) below of N-1 contravariant vectors becomes in this limit an ordinary covariant vector. If g=1 in x-space, then g' = RRT = 1 in x'-space and then that resulting vector can be considered either contravariant or covariant since both spaces are then Cartesian. This is the case with A = B x C under rotations in 3D space. On can think of the εabc as moving in this limit from a tensor density of weight -1 to an ordinary tensor. (c) Theorem about Totally Antisymmetric Tensors: there is really only one Theorem: Apart from a scalar factor, there exists only one totally antisymmetric (TA) tensor. Proof: Suppose there were two TA tensors called εabc... and rabc.... If two or more of the indices are equal, both tensors are 0, so for such index sets, one can say rabc.. = f εabc.. where f is any finite function whatsoever. Consider now the case where all the indices are distinct and therefore exhaust the set 123...N, and consider abc... to be a permutation of 123...N obtained by doing S pairwise swaps, abc... = P(123...) p = (-1)S If one were to associate a sign change with each swap, the total sign change would be p, the parity. Since ε and r are both TA tensors, each tensor can be "unwound" back to a standard index order by doing these S swaps, and the swaps will cause a total sign of p relative to that standard order, so rabc.. = p r123... // for example, r2134.. = (-1)1 r1234.. εabc.. = p e123... Define scalar function f ≡ r123... / e123... , whatever it might be. Then rabc.. = p(f e123...) εabc.. = p e123... and dividing these two equations one finds, rabc.. = f εabc.. which has now been shown valid for all index sets abc.. . Therefore, any "other" totally antisymmetric tensor is just a scalar function times the ε tensor. (d) The contravariant ε tensor Knowing nothing to start, assume that the famous εabc.. totally antisymmetric tensor transforms under F as a tensor density of some weight W which we hope to determine. Then ε'abc.. = J-W Raa' Rbb' ... εa'b'c'.. (*) Assume that εabc.. is the usual permutation tensor normalized to ε123...N = +1. This is the convention used by Weinberg p 99. This means each index swap changes the sign, and if two or more indices are the same, ε = 0. This is an important starting assumption, and from it most everything follows. Given this assumption, the RHS of (*) is totally antisymmetric (TA). The argument is given once here and then used later several times. Consider an a↔b swap. Then ε'bac.. = J-W Rba' Rab' ... εa'b'c'.. = J-W Rbb' Raa' ... εb'a'c'.. = J-W Raa' Rbb' ... (–εa'b'c'..) = – ε'abc.. The same result is true for any swap, thus RHS (*) = TA. Since according to section (b) there is only one TA tensor available, apart from a scalar function factor, it follows that ε'abc.. = Kεabc... where K is some scalar function, perhaps just a constant. Equation (*) above then reads K εabc... = J-W Raa' Rbb' ... εa'b'c'.. (**) Setting in the standard order, one finds that K ε123... = J-W R1a' R2b' ... εa'b'c'.. or K = J-W det(Rij) = J-W (J)-1 = J-(W+1) so now ε'abc.. = Kεabc... = J-(W+1) εabc... A second assumption is now made: that εabc.. (contravariant!) has the same value structure in any frame of reference, which is to say it is the same in x'-space as it is in x-space, ε'abc.. = εabc... This assumption is consistent with taking W = -1 in the previous equation. Again, this follows the convention of Weinberg p 99. Some authors instead arrange for the above equation to be true for the covariant ε tensors, and use then ε123...N = ε'123...N = +1, but we shall follow Weinberg. To summarize, assuming that εabc... is the usual permutation tensor normalized in the usual way, and assuming that ε'abc.. = εabc... so this tensor is the same in all frames or spaces, THEN one concludes that εabc... must transform as a rank-N tensor density of weight W = -1. That is to say, ε'abc.. = J Raa' Rbb' ... εa'b'c'.. This then is our second example of a tensor density. Viewed in this light, the tensor εabc.. is known as the Levi-Civita tensor. Tullio Levi-Civita (1873-1941). Italian, University of Padua 1892, with Ricci published the theory of tensor algebra in 1900 (see Refs.), which work assisted Einstein circa 1915 in formulating the theory of general relativity. The ε tensor bears his name. Sometimes the affine connection is called the Levi-Civita connection. (e) Some facts about the ε tensor 1. Consider ( based on section (b) 4 above), εabc... = gaa' gbb'..... εa'b'c'... This is again in the convention of Weinberg p 99 (4.4.10). Added sign s convention. Some authors make a special exception for the ε tensor and introduce an extra sign s into the above equation ( recall that s = -1 for special relativity) εabc... = s gaa' gbb'..... εa'b'c'... s = sign[det(gij)] Inserting such a sign renders any ε mixed tensor like εabc.. ambiguous when s =- 1, but is acceptable if one promises never to make use of a mixed ε tensor. We shall refer to these two methods as "the Weinberg convention" and the "added sign s convention", the former being assumed unless otherwise stated. Install the reference sequence to obtain ε123.. = g1a' g2b'..... εa'b'c'... = det(gij) = g // Section 5 (k) Similarly, ε'123.. = det(g'ij). To summarize, ε123.. = det(gij) = g // ε all-down index reference values ε'123.. = det(g'ij) = g' In the "added sign s convention", these last two equations would have sg = |g| and sg' = |g'| on the right which means then these two ε values would be always positive. 2. Take the same starting point as above εabc... = gaa' gbb'..... εa'b'c'... The RHS is a totally antisymmetric in indices abc... (see above) and can therefore be written RHS = C εabc... since we showed earlier that there is only one TA tensor apart from scalar C. Therefore εabc... = C εabc... (*) Insert the reference sequence ε123... = C ε123... = C But in 1 it was just showed that ε123... = det(gij) . Therefore C = det(gij) and then (*) says for the "Weinberg convention", εabc... = det(gij) εabc... = g εabc... // relating all down to all up ε'abc... = det(g'ij) ε'abc... = g' ε'abc... // = g' εabc... where the second line follows by the same argument. These equations relate all indices down to all up in the same space. Notice that both εabc... and ε'abc... are totally antisymmetric. In the "added sign s convention" the above equations are instead εabc... = |det(gij)| εabc... = |g| εabc... // relating all down to all up ε'abc... = |det(g'ij)| ε'abc... = |g'| ε'abc... // = |g'| εabc... 3. Divide the last two Weinberg convention equations to find that ε'abc... = [det(g'ij)/ det(gij)] εabc... = (g'/g) εabc... From Section 5 (k) this says ε'abc... = J2 εabc... and this same conclusion is valid for the "added sign s convention" as well since det(g) and det(g') always have the same sign as shown in Section 5 (k). Two comments: Although we set ε'abc... = εabc... by fiat, we cannot similarly set ε'abc... = εabc...by fiat. This latter result comes out being ε'abc... = J2 εabc... as just shown. The fact that ε'abc... = J2 εabc... does not say that εabc is a tensor density of weight -2 because there are no R factors showing. (See the section (a) definition of a tensor density transformation. ) (f) The covariant ε tensor According to section (b) 5, lowering indices does not change the weight of a tensor density. Section (d) showed that εabc.. is a tensor density of weight -1, so we know right away that εabc.. is also a tensor density of weight -1. Nevertheless, it is interesting to see what happens when the same method used in section (c) for eabc... is applied to eabc... . We start by assuming εabc.. is a tensor density of some unknown weight W, ε'abc.. = J-W [ Raa' Rbb'.... εa'b'c'..] (*) Section (e) 2 noted that εa'b'c'... is a totally antisymmetric tensor, and therefore as in section (d) one concludes that the RHS of (*) is also totally antisymmetric and can be written as RHS(*) = K εabc.. so (*) then says K εabc.. = J-W [ Raa' Rbb'.... εa'b'c'..] (**) Use section (e) 2 to set εa'b'c'.. = det(gij) εa'b'c'.. inside the bracket, K εabc.. = J-W [ Raa' Rbb'.... det(gij) εa'b'c'..] , and then install the reference sequence on both sides K ε123.. = J-W [ R1a' R2b'.... det(gij) εa'b'c'..] But section (e) 1 says ε123.. = det(gij), so cancel det(gij) on both sides to get K = J-W [ R1a' R2b'.... εa'b'c'..] = J-W det(Rij) = J-W det(Sji) = J-W J = J-(W-1) So here the result is K = J-(W-1) whereas in section (d) the result was K = J-(W+1) . In the current case, since (*) and (**) have the same RHS, setting the LHS's equal says ε'abc.. = K εabc.. = J-(W-1) εabc.. But section (e) 3 said that ε'abc.. = J2 εabc.. and therefore one gets W = -1. Thr conclusion is that εabc... transforms with weight -1, the same as εabc..., so (*) becomes ε'abc.. = J [ Raa' Rbb'.... εa'b'c'..] (g) Generalized cross products In Appendix A (c) the following cross product of N-1 vectors is considered (converted now to standard notation) Qa ≡ εabc...x BbCcDd.....Xx or Q = B x C x D .... x X If the vectors B,C,D..X are all contravariant vectors, then applying the rule of section (b) 3, one concludes that, since ε is a tensor density of weight -1 and since all the RHS vectors have weight 0, the object Qa is a covariant vector density of weight = -1, and thus has this transformation rule Q'a = J RabQb Similarly, one may consider Qa ≡ εabc...x BbCcDd.....Xx . If vectors B,C,D...X are covariant vectors, then Qa is a vector density of weight -1 and Q'a = J RabQb (h) The tensorial nature of curl B It has just been shown that C = A x B is a vector density of weight -1, this being a special case of the generalized cross product discussion above. As noted in section (a) 7, if R happens to be a rotation, then C is in fact a tensorial vector. One might conjecture that C = x B is also a vector density of weight -1, and that conjecture is correct as is now shown. Consider Cn = εnab ∂aBb where B is assumed to be an tensorial vector. It is helpful to write this equation in the following manner Cn = εnab [ ∂aBb – ∂bBa ]/2 where the second term is the same as the first term, since – εnab ∂bBa = εnba ∂bBa = εnab ∂aBb . Recall now from Section 7 (v) that the covariant derivative of a vector is given by Bb;a = ∂aBb – Γcab Bc where the affine connection Γcab is symmetric under a↔b. Therefore Bb;a – Ba;b = [∂aBb – Γcab Bc] - [∂bBa – Γcba Bc] = ∂aBb – ∂bBa Therefore Cn can be expressed as Cn = εnab [Bb;a – Ba;b ]/2 so by the same ε anti-symmetry noted above the final result is Cn = εnab Bb;a The major feature of Bb;a -- as discussed in Section 7 (v) -- is that it is a rank-2 tensor if B is a tensorial vector. The weight addition rule of section (b) 3 can then be applied to the RHS above. Since ε has weight -1 and B has weight 0, the conclusion is that Cn is a vector density of weight -1. Thus, C = curl B is a vector density of weight -1. If the underlying transformation F is a rotation, C becomes an ordinary vector as per section (a) 7. (i) Tensor E as a weight 0 version of ε : three conventions Equations in the "Weinberg Convention" In this section it is assumed that gij and gij raise and lower indices of the ε tensor just as they do for any other tensor (Weinberg convention). In sections (d) and (e) it was established that εabc... = g εabc... ε123.. = +1 ε123... = g ε'abc... = g'ε'abc... ε'123.. = +1 ε'123... = g' ε'abc... = J2 εabc... = (g'/g) εabc... ε is a rank-N tensor of weight W = -1 Again, just in passing, notice that ε'abc...= J2 εabc...does not say ε has weight -2 because the R factors are not present on the right side. Consider now the following new objects defined by (sg = |g|, s= sign(g) = sign(g') as in Section 5 (k)) Eabc... ≡ |g|-1/2 εabc... => E123... = |g| -1/2 g = |g| -1/2 s |g| = s |g|1/2 E'abc... ≡ |g'|-1/2 ε'abc.. . => E'123... = |g'| -1/2 g' = |g'| -1/2 s |g'| = s |g'|1/2 It was shown in Section 5 (k) that g' = J2g so that g transforms as a scalar density of weight -2. Since the sign of g and g' are the same, if follows that (sg) = |g| is also a scalar density of weight -2, and then the quantity |g|-1/2 transforms as a scalar density of weight +1, since |g'|-1/2 = J-1 |g|-1/2. Looking at the equation Eabc... ≡ |g|-1/2 εabc... above, and using the weight summation rule of section (b) 3, one concludes at that Eabc... transforms as a tensor of weight (+1) + (-1) = 0, and so Eabc... is an ordinary tensor. This is the motivation of the above definitions. It was shown at the end of Section 7 (u) that a tensor density equation with matching weights is "covariant", so one is not surprised to see the second line above being the same as the first line but everything is primed. Raising indices on both sides gives Eabc... ≡ |g|-1/2 εabc... => E123... = |g|-1/2 E'abc... ≡ |g'|-1/2 ε'abc... => E'123... = |g'|-1/2 To compare Eabc... and Eabc... , Eabc... = |g|-1/2 εabc... = |g|-1/2 g εabc... = s |g|-1/2 |g| εabc... = s |g|1/2 εabc... Eabc... = |g|-1/2 εabc... so that, Eabc... = s|g| Eabc... = g Eabc... Summarizing, E123... = |g|-1/2 E123... = s|g|+1/2 Eabc... = g Eabc... = s|g| Eabc... E'123... = |g'|-1/2 E'123... = s|g'|+1/2 Eabc... EABC... = |g|-1 εabc... εABC... E'abc... E'ABC... = |g'|-1 ε'abc... ε'ABC... Since |g|-1 is a scalar density of weight +2 and each ε has weight -1 and each E has weight 0, one is happy to see the weights balance of the two sides of this pair of covariant equations. Equations in the "added sign s convention" The previous section shows how things work out using the "Weinberg convention" noted at the start of section (e). Here is the previous section redone in the "added sign s convention": In section (e) it was established that ( g = det(gij)) εabc... = sgεabc... ε123.. = +1 ε123... = sg = |g| // g → sg ε'abc... = sg'ε'abc... ε'123.. = +1 ε'123... = sg' = |g'| // g' → sg' ε'abc... = |J|2 εabc... = (g'/g) εabc... ε is rank-N tensor of weight W = -1 // same Consider the following new objects defined by ( sg = |g|, s= sign(g) = sign(g') as in Section 5 (k) ) Eabc... ≡ |g| -1/2 εabc... => E123... = |g| -1/2 |g| = |g|1/2 E'abc... ≡ |g'| -1/2 ε'abc... => E'123... = |g'|-1/2 |g'| = |g'|1/2 But the same argument given above, Eabc... is an ordinary covariant tensor (ie, weight = 0). However, the indices cannot be raised by gij. In this convention then one must make independent definitions of the contravariant components as follows, Eabc... ≡ |g|-1/2 εabc... => E123... = |g|-1/2 E'abc... ≡ |g'|-1/2 ε'abc... => E'123... = |g'|-1/2 To compare Eabc... and Eabc... , Eabc... = |g|-1/2 εabc... = |g|-1/2 |g| εabc... = |g|-1/2 |g| εabc... Eabc... = |g|-1/2 εabc... so that Eabc... = |g| Eabc... Summarizing, E123... = |g|-1/2 E123... = |g|+1/2 Eabc... = |g| Eabc... E'123... = |g'|-1/2 E'123... = |g'|+1/2 E'abc... = |g'| E'abc... Eabc... EABC... = |g|-1 εabc... εABC... E'abc... E'ABC... = |g'|-1 ε'abc... ε'ABC... In this "added sign s" convention, all these summarized results involve only |g| and there are no factors of s floating around. The cost of this benefit is a lack of true covariance (when s=-1), as demonstrated in section (k) below. Equations in the "Ricci-Levi-Civita convention" Ricci and Levi-Civita use the "added s convention" but add a factor σ = sign(det(S)) into their definition of E ( see their paper p 135 or Hermann pp 31-21) so that Eabc... ≡ σ|g| -1/2 εabc... => E123... = σ|g| -1/2 |g| = σ|g|1/2 E'abc... ≡ σ|g'| -1/2 ε'abc... => E'123... = σ|g'|-1/2 |g'| = σ|g'|1/2 Eabc... ≡ σ|g|-1/2 εabc.. => E123... = σ|g|-1/2 E'abc... ≡ σ|g'|-1/2 ε'abc... => E'123... = σ|g'|-1/2 Summarizing, E123... = σ|g|-1/2 E123... = σ|g|+1/2 Eabc... = |g| Eabc... E'123... = σ|g'|-1/2 E'123... = σ|g'|+1/2 E'abc... = |g'| E'abc... Eabc... EABC... = |g|-1 εabc... εABC... E'abc... E'ABC... = |g'|-1 ε'abc... ε'ABC... Notice that in all three conventions, last equation pair is the same. Since Ricci and Levi-Civita did not raise and lower individual indices in their 1900 paper, they were not concerned about their convention being non-covariant in that sense. (j) Representation of ε, εε and contracted εε as determinants 1. Theorem about a certain permutation sum Consider the following object Q defined as a signed permutation sum of the product of N matrix elements of a matrix Mij, Qabc..x ≡ ΣP p P2(Ma1Mb2 Mc3.....MxN) In this equation, P2 represents a permutation of the set of 2nd indices of the N matrix elements, and the sum is over all N! such permutations. There are many ways to arrive at a given permutation of 123...N by doing pairwise swaps, but for all these ways, the number of swaps S will be either even or odd. The parity p of a permutation is defined then as (-1)S and this p appears in the above sum. If one were to swap 2↔3 on the right above, each permutation would have S→ S+1 since an extra swap is needed to undo 2↔3. Thus, all parities p → -p and in fact the whole object negates. But the swap 2↔3 is the same as b↔c since Mb3 Mc2 = Mc2 Mb3. Applying this argument to any pair of indices, one concludes that Qabc..x is totally antisymmetric and therefore can be written as K εabc...x : ΣP p P2(Ma1Mb2 Mc3.....MxN) = K εabc...x Setting abc..x to 123..N, one gets. ΣP p P2(M11M22 M33.....MNN) = K The left side of this last equation can be written as ΣP p P2(M11M22 M33.....MNN) = Σabc..x εabc...x M1aM2b M3c.....MNx because p = εabc...x correctly assesses the parity of any given permutation. But this object is simply det(M) so the conclusion is that K = det(M) and then ΣP p P2(Ma1Mb2 Mc3.....MxN) = det(M) εabc...x Consider now the following matrix where abc..x is some permutation of 123...x, where M(abc..) = Ma1 Ma2 Ma3 ... MaN Mb1 Mb2 Mb3 ... MbN Mc1 Mc2 Mc3 ... McN ... Mx1 Mx2 Mx3 ... MxN By rearranging the rows into their normal numerical order, one obtains matrix M, but incurs a sign from the various row swaps which sign is just εabc..x. Therefore det(M(abc..) ) = εabc...x det(M) and therefore ΣP p P2(Ma1Mb2 Mc3.....MxN) = det(M(abc..) ) = det(M) εabc...x The permutation sum is thus just the determinant of matrix M(abc..). The first term in the permutation sum, the term with an identity permutation, corresponds to the product of the diagonals of that matrix. 2. Application of the theorem to M = δ: a representation of ε Apply the above theorem to matrix M = 1 ≡ δ, the identity matrix, having Mij = δi,j. Clearly det(δ) = 1 and one then has ΣP p P2(δa,1δb,2 δc,3.....δx,N) = det[δ(abc..) ] = εabc...x Thus is obtained the famous representation of εabc..x as a certain determinant of Kronecker deltas, εabc...x = det[δ(abc..) ] where δ(abc..) = δa,1 δa,2 δa,3 ... δa,N = Ra δb,1 δb,2 δb,3 ... δb,N = Rb δc,1 δc,2 δc,3 ... δc,N = Rc ... δx,1 δx,2 δx,3 ... δx,N = Rx where, for future use, each row vector has been given a name like Ra where (Ra)i = δa,i The conclusion then is that which is the same as εabc...x = ΣP p P2(δa,1δb,2 δc,3.....δx,N) . 3. Outer product of two ε tensors. Consider now εabc...x = det and εa'b'c'...x' = det Then εabc...x εa'b'c'...x' = det det = det det ( Ra' Rb' ... Rx') = det { ( Ra' Rb' ... Rx') } which is the determinant of this matrix Ra Ra' Ra Rb' Ra Rc' ... Ra Rx' Rb Ra' Rb Rb' Rb Rc' ... Rb Rx' Rc Ra' Rc Rb' Rc Rc' ... Rc Rx' ... Rx Ra' Rx Rb' Rx Rc' ... Rx Rx' A typical element of this matrix is given by Rc Rb' = (Rc)i(Rb')i = δc,i δb',i = δc,b' so that matrix can be written as δa,a' δa,b' δa,c' .... δa,x' δb,a' δb,b' δb,c' .... δb,x' δc,a' δc,b' δc,c' .... δc,x' ≡ δ(abc..x; a'b'c'..x') .... δx,a' δx,b' δx,c' .... δx,x' where we have made up a name for this matrix as shown. The conclusion then is that which is the same as εabc...x εa'b'c'...x' = ΣP p P2(δa,a'δb,b'δc,c'.....δx,x') As usual, the argument of P2 is the product of the diagonal elements of the matrix. 4. Contracting the first index of the outer product of two ε tensors. Consider what happens if one sums on the first index of the εε product: Σa εabc...x εab'c'...x' For fixed given values of bc..x and b'c'...x' , there is only one way this sum can be non-zero. In that one way, bc..x and b'c'...x' must each be permutations of the set {12..N exclude A} where A is the "hit value" of a in the sum on a. Then Σa εabc...x εab'c'...x' = εAbc...x εAb'c'...x' = ΣP p P2(δA,Aδb,b'δc,c'.....δx,x') = ΣP p P2(δb,b'δc,c'.....δx,x') (*) where in this last expression the sum can be regarded as being over permutations where b'c'...x' is a permutation of b,c..x. Each of these lists of integers is in turn a permutation of {12..N exclude A}. Now, parity p = (-1)S where S is a number of swaps it takes to connect b'c'...x' with b,c..x, since a = a' = A. One might wonder if the overall sign of the RHS of the last equation is correct. A check of the first term in this sum which is just δb,b'δc,c'.....δx,x' shows that this overall sign is indeed correct. This first term must be positive because the product of two ε's is either +1 or 0. As an example, Σa εabc εab'c' = ΣP p P2(δb,b'δc,c') = δb,b'δc,c' – δb,c'δc,b' The permutation sum shown on the right side of (*) is the determinant of δ(abc..x; a'b'c'..x') but with the first row and column crossed out. It can then be thought of as either the minor or cofactor of the element aa of this big δ matrix. Therefore, Σa εabc...x εab'c'...x' = [cof δ(abc..x; a'b'c'..x')]aa where the notation cofM refers to a matrix of cofactors with elements [cofM]ij. Don't confuse the a on the right side with the local dummy summation index a on the left side. The conclusion then is that (implied summation on a on the LHS) 5. Contracting two or more indices of the outer product of two ε tensors. Consider what happens if one sums on the first two indices of the εε product: Σa,b εabcd...x εabc'd'...x' For fixed given values of c,d..x and c'd'...x' , in order for this double sum to be non-zero, the index sets cd..x and c'd'...x' must each be permutations of the set {12..N exclude A,B} where A,B are a pair of hit values for the a and b sums. If a=A and b=B is hit value, then so is a=B and a=A, so there are 2! contributing terms in the sum, and each term is +1. Therefore Σa,b εabcd...x εabc'd'...x' = 2! εABc...x εABc'...x' = 2! ΣP p P2(δA,AδB,B'δc,c'δd,d'.....δx,x') = 2!ΣP p P2(δc,c' δd,d'.....δx,x') (*) where in this last expression the sum is over permutations where c'd'...x' is a permutation of c,d....x. Each of these lists of integers is in turn a permutation of {12..N exclude A,B}. Now parity p = (-1)S where S is a number of swaps it takes to connect c'd'...x' with c,d....x. Since the product of two ε's is either +1 or 0, the overall sign of the right side shown must be correct. As an example, Σa,b εabc εabc' = 2!ΣP p P2(δc,c') = 2 δc,c' If c = c' = 2, then this says Σa,b εab2 εab2 = ε132 ε132 + ε312 ε312 = 1 + 1 = 2 The permutation sum shown on the right side of (*) is the determinant of δ(abc..x; a'b'c'..x') but with the first 2 rows and columns crossed out. Therefore, Σa,b εabcd...x εabc'd'...x' = 2! { [cof δ(abc..x; a'b'c'..x')]aa}bb The conclusion then is that (implied summation on a,b on the LHS) This pattern continues as more indices are contracted. If three indices a,b,c are contracted, there will then be 3! hit values which are A,B,C and its permutations, and one just repeats the above discussion. The result will then be Σa,b,c εabcd...x εabcd'...x' = 3! {{ [cof δ(abc..x; a'b'c'..x')]aa}bb}cc The conclusion then is that (implied summation on a,b,c on the LHS) Eventually one arrives at a point where all but one of the indices are summed, so that εabcd...x εabcd...x' = (N-1)! |δx,x'| = (N-1)! δx,x' an example being εabc2 εabc2 = 3! δ22 = 3! = ε1342 ε1342 + ε1432 ε1432 + 4 more terms = 1+1+4 = 6 The final point is that at which all indices are summed, with result εabcd...x εabcd...x = N! and example of which is εabcεabc = ε1232 + ε2132 + 4 more terms = 1 + 1 + 4 = 6 6. Summary of Results εabcd...x εabcd...x' = (N-1)! δx,x' εabcd...x εabcd...x = N! (k) Covariant forms of the previous section results The results above were all developed in Cartesian x-space where up and down indices on the ε's did not matter. The rules for converting any result above to covariant form are as follows: Weinberg convention: write the left side as either |g|-1 ε***** ε***** or as E***** E***** .The objects with indices as shown by asterisks are true tensors (weight 0). write the right side replacing every δa,b by δab → gab as shown in Section 7 (m). Then the right side will also be a true tensor. Example: The εε product for N=2 with no indices summed was written above as (g = 1) εabεa'b' = = δa,a' δb,b' – δa,b' δb,a' The covariant form is as follows, where now g is some arbitrary metric tensor for x-space, EabEa'b' = |g|-1 εabεa'b' = = gaa' gbb' – gab' gba' The equation in x'-space would then be E'abE'a'b' = |g'|-1 ε'abε'a'b' = = g'ab g'ab' – g'a'b g'a'b' because true tensor equations are "covariant"(Section 7 (u)). One can raise and lower individual indices to get for example these valid tensor equations which are 3 members of the family of 4! = 24 tensor equations obtained by raising and lowering indices: EabEa'b' = |g|-1 εabεa'b' = = gaa' gbb' – gab' gba' EabEa'b' = |g|-1 εabεa'b' = = gaa' gbb' – gab' gba' EabEa'b' = |g|-1 εabεa'b' = = gaa' gbb' – gab' gba' and of course in x'-space the equations are the same but everything is primed. Added-sign-s and Ricci-Levi-Civita conventions: Do the above two bullet items, then add an overall sign s to the right side, because εabc.. = s g εabc... in these conventions instead of εabc.. = g εabc... so that εabc..(Weinberg) = sεabc..(added-sign). Example: The first equation above becomes ( εa'b'→ s εa'b') EabEa'b' = |g|-1 εabεa'b' = s = s ( gaa' gbb' – gab' gba') The second equation is undefined (when s=-1), and the third equation is EabEa'b' = |g|-1 εabεa'b' = = gaa' gbb' – gab' gba' The first and third equations are true tensor equations, except individual indices cannot be raised and lowered. If one were doing some significant work involving covariance and s=-1, it would certainly seem advisable to use the Weinberg convention since it is completely "covariant" for either sign of s. Appendix E: Tensor Expansions: direct product, polyadic and operator notation This entire section uses the general Picture A context where x-space need not be Cartesian, The Standard Notation is used throughout. (a) Direct Product Notation The key tool required for the expression of tensor expansions is the notion of a direct product of n tensorial vectors defined in this simple way, (ABC ...)abc... ≡ AaBbCc..... (ABC ...)abc... ≡ AaBbCc..... etc The tensor ABC ... is nothing more than the outer product of vectors A,B,C as discussed in Section 7 (a) for contravariant vectors, but later extended to any mixture of vector types. As noted in Section 7 (j), one can define a direct product of two rank-2 tensors in this way, (MN)ab,AB ≡ MaANbB // rank = n = 2; number of tensors = I = 2 (MN)ab,AB ≡ MaANbB etc and then the same idea can be applied to form a direct product of tensors of any rank, for example (MN)ab,AB,αβ = MaAαNbBβ etc // rank = n = 3; number of tensors = I = 2 On the left side the number of groups of indices equals the tensor rank n, and the number of indices within each group matches the number I of tensors being direct-product-multiplied. In what follows, only the direct product of vectors shall be considered. One can define the dot product of two direct-product-space vectors in this obvious manner, (ABC ...) (A'B'C' ...) ≡ (ABC ...)abc... (A'B'C' ...)abc = AaBbCc..... A'aB'bC'c..... = AA' BB' CC' ... where of course the indices abc can be tilted in any way desired according to Section 7 (k). (b) Tensor Expansions and Bases Let bi be an arbitrary complete set of basis vectors in x-space. As shown in Section 6 (b) there exists a unique set of dual basis vectors bi (also in x-space) such that bi bj = δij. Consider then the following expansion of a rank-3 tensor A A = Σijk αijk (bibjbk) where (bibjbk ...)abc = (bi)a (bj)b (bk)c The coefficients αijk can be obtained by dotting both sides with (bi'bj'bk') and using (bi'bj'bk') (bibjbk) = bi' bi bj' bj bk' bk = δi'iδj'jδk'k . The result is then αijk = A (bibjbk) = Aabc (bibjbk)abc = Aabc (bi)a (bj)b (bk)c . (*) where Aabc are the contravariant components of tensor A in x-space, and (bi)a are the covariant components of vector bi in x-space. In this manner, a tensor A of any rank can be expanded on an arbitrary complete set of basis vectors, and the coefficients of that expansion can be obtained by the inversion shown above for rank 3. Two special bases are of interest. The ui are the axis-aligned basis vectors in x-space as discussed in see Section 7 (s). For these basis vectors, one has (ui)a = δia and (ui)a = δia . If one considers this expansion, A = Σijk αijk (uiujuk) where (uiujuk ...)abc = (ui)a (uj)b (uk)c = δia δjb δkc then the coefficients are found to be αijk = A (uiujuk) = Aabc δia δjb δkc = Aijk so the coefficients are exactly the x-space contravariant components of the tensor A. Thus A = Σijk Aijk (uiujuk) On the other hand, if ei are the tangent base vectors in x-space (see Sections 3), the dual vectors are the ei and from Section 7 (s) one has (ei)a = Sai = Ria and (ei)a = Sai = Ria . If one considers the expansion A = Σijk αijk (eiejek) where (eiejek ...)abc = (ei)a (ej)b (ek)c = Ria Rjb Rkc then the coefficients are found to be αijk = A (eiejek) = Aabc (ei)a(ej)b(ek)c = Aabc Ria Rjb Rkc = Ria Rjb Rkc Aabc = A'ijk and thus the coefficients in this case are exactly the x'-space contravariant components of tensor A, as shown in Section 7 (j). Thus, A = Σijk A'ijk (eiejek) Expansions like the above are the generalizations to tensors of any rank of these vector expansions stated in Section 7 (s), A = Σiαi bi αi = bi A // arbitrary basis A = ΣiAi ui // axis aligned unit vectors A = ΣiA'i ei // tangent base vectors where we continue to write rank-1 tensors in bold font: A. To summarize, here is the general rank-n tensor expansion for an arbitrary basis, and then for the two specific bases just discussed: A = Σijk... αijk... (bibjbk...) αijk... = Aabc... (bi)a (bj)b (bk)c... A = Σijk... Aijk... (uiujuk...) Aijk... = contravariant components of A in x-space A = Σijk... A'ijk... (eiejek...) A'ijk... = contravariant components of A in x'-space Orthonormal basis: If the basis vectors bi happen to be orthonormal as defined by bi bj = δij then bi = bi because the dual basis is unique. As indicated in (*) above, this implies that coefficient αijk is unchanged if any or all indices are lowered, as if these αijk were components of a tensor in some Cartesian space. That Cartesian space is in fact the x'-space that would arise if transformation F were custom-selected such that the bi were the tangent base vectors ei for that F, for then g'ij = ei ej = δi,j so that x'-space would in fact be Cartesian. But for a pre-determined F, the αijk are just some coefficients and are not components of a tensor relative to F, and it just happens that αijk = αijk etc. An example of orthonormal basis vectors arises if bi = i ≡ ei/h'i and x'-space has a diagonal metric tensor g'ab = h'a2δa,b. One then has i j = δi,j since i j = ei ej / (h'i h'j) = g'ij/ (h'i h'j) = h'i2δij/ (h'i h'j) = δi,j . Then since the dual basis is unique, one has i = i and then αijk(any up/down) = Aabc (i)a (j)b (k)c = A'ijk (h'ih'jh'k) where the last expression comes from the third expansion shown above. Tensor density. If A is a tensor density of weight W, the general rule is to make this replacement: A'ijk... → J WA'ijk... so the third general expansion above would be written A = J W Σijk... A'ijk... (eiejek...) A'ijk... = contravariant components of A in x'-space As justification for this rule, start with a regular tensor transformation for A, A'ijk... = Rii' Rjj' Rkk'..... Ai'j'k'... The rule gives J W A'ijk... = Rii' Rjj' Rkk'..... Ai'j'k'... or A'ijk... = J-W Rii' Rjj' Rkk'..... Ai'j'k'... which is the correct form for the transformation of a tensor density (Appendix D). Expansion of tensor-like objects. If Aijk is some "tensor like" object having three indices (such as ∂iTjk) one can still do the three expansions shown above but the results would have to be restated this way: A = Σijk... αijk... (bibjbk...) αijk... = Aabc... (bi)a (bj)b (bk)c... A = Σijk... Aijk... (uiujuk...) Aijk... = components of A in x-space A = Σijk... Aijk... (eiejek...) Aijk... = Ria Rjb Rkc Aabc Since A is not a tensor, in this case Aijk... are not the contravariant components of tensor A in x'-space. (c) Polyadic Notation Some fields of study historically use "polyadic notation" as follows, (ABC...) ≡ ABC ... It is sometimes a bit disturbing to modern readers to see bolded vectors stacked directly against each other, but the direct product makes the meaning clear. For arbitrary basis vectors, one would then have, for example, (bibjbk...) ≡ bibjbk ... Sometimes this basis vector notation is compressed even more, to wit, i j k ... ≡ (bibjbk...) ≡ bibjbk ... although this notation seems to be mostly used when the bi are the unit vectors ui. In all these notations, one must be aware that the symbols do not "commute". For example i j = (bibj) = bibj => (i j )nm = (bibj)nm = (bibj)nm = (bi)n (bj)m (j i )nm = (bjbi)nm = (bjbi)nm = (bj)n (bi)m ≠ (i j )nm and therefore one cannot write i j = j i. The general expansion stated above now appears as A = Σijk... αijk... (bibjbk...) = Σijk... αijk... (bibjbk...) = Σijk... αijk... (i j k ...) where αijk... = A (bibjbk...) = A (bibjbk...) = A (id jd kd ... ) = Aabc... (bi)a (bj)b (bk)c... where we have just made up a notation id to stand for the dual vector bi. One can find further discussion of polyadic notation for example in Backus. (d) Dyadic Products When two vectors A and B are combined in polyadic notation, the result is called a dyadic product (AB) [ also known as a dyad or just a dyadic ] (AB)ij ≡ AiBj // = (AB)ij In this notation, the expansion given above for a rank-2 tensor becomes A = Σij αij (bibj) αij = Aab(bi)a (bj)b = Aab (bibj)ab Notice that the dyadic product (AB) is a rank-2 tensor if we assume that the underlying Ai and Bi are the x-space contravariant components of tensorial vectors A and B (which we normally assume). As a reminder, x-space need not be Cartesian. In section (g) below it will be shown that the matrix (AB)ij can be associated with an operator (AB) in the un basis so (AB)ij = <ui |(AB)| uj >, but this interpretation is not necessary for what follows. (e) Transpose notation for dyadics Superscript T as usual indicates the transpose of a vector or matrix. For a vector V, certainly Vi = (VT)i, meaning the object in the ith column of V is the same as the object in the ith row of VT . Therefore one can express the dyadic product in this more down-to-earth manner, (ab)ij ≡ aibj = ai (bT)j = (abT)ij or ab = abT Here one knows that ab is a "dyadic" because there is no other meaning for two bolded column vectors abutting each other with no intervening operator, so no special notation like [ab] is needed to indicate that ab is a dyadic. The object abT on the other hand has a well-defined meaning in matrix algebra, abT = (b1 b2) = = a matrix and one sees that in fact (ab)ij = (abT)ij = aibTj = aibj . Meanwhile, the object aTb is just a number, aTb = a b = (a1 a2) = a1b1 + a2b2 = a scalar (if a and b are vectors) This transpose notation can then be applied to the dyadic expansion of a 2x2 matrix A, A = Σij αij bibj = Σij αij bibjT = α11 b1 b1T + α12 b1 b2T ... In the special case that the bi are the unit vectors ui , and assuming N = 2 dimensions, one has A = Σnm Anm unum = Σnm Anm unumT = A11 u1 u1T + A12 u1 u2T + A21 u2 u1T + A22 u2 u2T = a matrix with A12 in the upper right corner where un is a column unit vector and unT is the corresponding row unit vector (see comments in Section 3 (c) about "unit" vectors). For example, u1u2T= ( 0 1) = Obviously this matrix visualization is valid for any dimension N, not just N=2. For rank n > 2, however, this transpose-of-vector concept does not conveniently generalize. For n=3 the object uaubuc would be a cube of zeros with a single 1 located at coordinates a,b,c, and so on for n > 3. One cannot write this as uaubucT for example. The direct product or polyadic notation seems clearest for rank n > 2. (f) Large and small dots used with dyadics Sometimes a small-size dot • is used to indicate the action of a dyadic (matrix) on a vector. If A is a dyadic then one defines: A• c ≡ Ac = a column vector => (A• c)i = (Ac)i = Aijcj => A• c = ΣijAijcj ui c • A ≡ cTA = a row vector => (c • A)i = (cTA)i = cjAji => c • A = ΣijcjAji ui d • A • c = dTAc = a number = diAijcj It then follows that, for the particular dyadic A = ab , (ab) • c ≡ (ab) c = (abT)c = a(bTc) = a ( b c) = ( b c) a = a column vector c • (ab) ≡ cT (ab) = cT(abT) = (cTa) bT = (c a) bT = a row vector d • (ab) • c = dT (ab) c = dTabT c = (dTa)( bT c) = (d a)(b c) = a number Here is more detail on the first line of the above group showing a skeletal matrix structure, (ab) c = (a bT)c = abTc = a(bTc) = a(bc) { } = {(b1 b2)} = (b1 b2) = { (b1 b2) } = bc The same small dot is used to indicate the product of two dyadics, which is to say, matrix multiplication A• B ≡ AB Regarding this small size dot • : (1) from a matrix algebra point of view, it is completely superfluous except in the case c • A ≡ cTA ; (2) it is completely different from the dot used in bTc = bc . It is this larger dot which was the subject of Section 6 (b); (3) The next section provides an explanation of the small dot as part of an operator interpretation for dyadics. (g) Operators and Matrices for Rank-2 tensors Operator concept. As discussed in Section 5 (i), x-space and x'-space of Picture A are both N-dimensional real Hilbert Spaces with scalar product indicated by the large dot , and one can regard V as a vector in either space. Expressed as a "vector" in x-space one can write, as done above with generic basis bi , V = Σi αi bi Moreover, one can regard a rank-2 tensor A as an "operator" in this Hilbert space, A = Σij αij bibjT Application of (bT)n on the left and bm on the right, and then a double use of (bT)nbi = bn bi = δni gives αnm = (bT)n A bm Here, one regards A as an operator in the x Hilbert space, whereas αnm is a "matrix" which is associated with the operator A in the particular bn basis. The idea of A as operator has an abstract meaning distinct from the matrix Aij. In the above equation the symbol A is this abstract operator and (bT)n A bm has a meaning distinct from our interpretation of it in terms of the matrix combination of three objects. In the matrix interpretation, one writes (bT)n A bm = [(bT)n]i Aij [bm]j = [bT]i Aij [bm]j and only then does A become a "matrix". This matrix happens to be the contravariant Aij matrix because we happened to select the un basis to write the components like [bm]j = uj bm. Bra-ket Notation. For the author of this document, the bra-ket notation commonly used in quantum mechanics (Paul Dirac 1939) provides a clean way to look at a rank-2 tensor A as an operator. It is true that in quantum mechanics one usually deals with infinite dimensional Hilbert spaces and complex numbers, but the formalism applies just as well to real Hilbert spaces with finite dimensions. In bra-ket notation one writes bi → |bi>, biT→ <bi| , so that the above equations become |V> = Σi αi |bi> <bj|bi> = δji orthogonality of the basis A = Σij αij | bi> <bj| 1 = Σi | bi><bi| completeness of the basis αnm = <bn | A | bm > <U | V> = U V = scalar product In this notation, the N |bi> are a set of basis vectors which span an N-dimensional real Hilbert Space, while <bi| span the so-called adjoint (or transpose in our case) Hilbert Space. One then refers to αnm as the "matrix element of the operator A in the bi basis ". In this notation, based on what was presented earlier, one can write, Anm = <un | A | um > = the x-space components of tensor A A'nm = <en | A | em > = the x'-space components of tensor A In all these equations the covariant indices can be moved up and down in the usual manner. Notice in the last two lines that the operator A between the vertical bars is the exact same operator on both lines. The matrices are different not because the operator has changed, but because the basis vectors are different. The point here is that one really can regard a rank-2 tensor A as an operator and not a specific matrix. In a given basis, that operator "has" a certain matrix. One might write the associated matrix as [A(b)]nm = <bn | A | bm >. That is to say, the matrix needs some label like (b) to indicate the basis used to define the matrix. In our notation, Anm with no label refers to the matrix associated with the um basis, while A'nm goes with the em basis. Bases are related by a transformation. Consider again ( note that |i> ≡ |ui> on the next line ) [A(b)]nm = <bn | A | bm > = (bn)T A bm = [bn]i Aij [bm]j // = <bn|i><i|A|j><j|bm> where the subscripts i and j are those associated with the un basis. Defining Bjm ≡ [bm]j one has [A(b)]nm = (BT)ni Aij Bjm or [A(b)]nm = (BT)ni Aij Bjm // lower m and reverse the tilt of index j or A(b) = BT A B // tilted matrix mulitplication as per Section 7 (i) which shows that the matrix elements [A(b)]nm are related to the Aij by a "congruence transformation" with a matrix B whose columns are the basis vectors bm . When bm = um , matrix B is the identity matrix, and when bm = em , one has Bjm ≡ [bm]j = (em)j = Rmj = Sjm . Thus B = S in this case. It was shown in Section 7 (i) that in standard notation S is real orthogonal, so in fact one has for the bm = em basis, A(e) = BT A B = ST A S = S-1 A S = R A R-1 and our two matrices are related by similarity by our old friends S and R as shown. More on bra-ket notation and its relation to the small dyadic dot. Consider the following facts, <d | A | c > = dT A cm = dT [A c ] = <d |Ac > <d | A | c > = dT A cm = [ dT A] c = [AT d]T c = <ATd | c> where |(Ac) > = a new Hilbert space vector which results when A is applied to |c> <(ATd) | = a new transpose Hilbert space vector which results when A is applied to <d| So one has this general idea that <d | A | c > = <d |Ac > = <ATd | c> A | c > = |(Ac)> <d | A = <(ATd) | In this last line, the isolated A's are the same operator A sitting in the Hilbert space. This operator can "act" either to the right or to the left as shown. The object |(Ac)> ≡ |e> is some different vector in the Hilbert space (different from |c>), call it |e>, and the grouping (Ac) labels this vector. Similarly, <(ATd) | is vector <e| in the transpose Hilbert space. The distinction between A as an abstract operator in the Hilbert space, and the A in (Ac) and (ATd) = (dTA)T as vectors in the Hilbert space is a subtle one. It is just this distinction that is implied by the small dot in the dyadic notation discussed in the previous section, and here is the correspondence between the dyadic notation and the bra-ket notation: A• c = Ac d • A = (ATd)T = dTA d • A • c = dTAc A• B c A | c > = |Ac> <d | A = <(ATd) | <d | A | c > = <d | Ac > AB| c > In the rightmost column operator B is applied first to | c> to get vector |(Bc)>, and then operator A is applied to |(Bc)> to give yet another vector | (Abc)>. In bra-ket notation the product of two abstract operators is given just as AB, but in dyadic notation it is written A• B. Dyadics as operators. According to the above discussion, one can regard a dyadic (AB), being a rank-2 tensor, as an operator and not as a matrix. The matrix Tnm = (AB)nm = AnBm is specific to the un basis in x-space (again, one might have g ≠1) Tnm = (AB)nm = <un |(AB)| um > = (un)T A BT um = [(un)T]a Aa (BT)b [um]b = δna Aa Bb δmb = AnBm . In the generic bn basis one might then write something like [(AB)(b)]nm = <bn |(AB)| bm > It is to emphasize this operator view of a dyadic that Morse and Feshbach use fancy letters like U to represent dyadics. Then the small-dot notation U • B emphasizes the idea of an operator acting on a vector, equivalent to U| B> . Here then are a few quotes from Morse and Feshbach (an = un) to illustrate some of the notation described above. These authors are working in Cartesian space (g=1) where up and down indices don't matter. ( The first item here is our A• c = ΣijAijcj ui from above. ) Notice the impressive name "idemfactor" for the identity operator 1 = Σi | ai><ai| = Σi aiaiT = Σiaiai. Appendix F: The Affine Connection Γcab and Covariant Derivatives (a) Definition of Γ Context can be confusing in a discussion of the affine connection Γ, so we start with a modified Picture C in which the quasi-Cartesian space on the right is called ξ-space instead of x(0)-space as in Picture C. The notation ξi for the coordinates of ξ-space seems traditional in general relativity writing, an application in which the Γ object appears frequently. Recall from Section 1 that the metric tensor G is a diagonal matrix whose elements are independently +1 or -1. And recall from Section 5 (b) that in x-space, gab = RaiRbjGij. If G = 1, then the xi coordinates of x-space are the "curvilinear coordinates" and the ξi are the "Cartesian coordinates". In Picture C1 the tangent base vectors en exist in ξ-space and the components of en are given by (en)i = Rni, as in Example 1 of Section 3 (polar coordinates). Since matrix R is the linearization of transformation F at ξ and since x = F(ξ), one can regard R as a function either of ξ or x. We choose x as the variable and write [en(x)]i = Rni(x). For example, Example 2 of Section 3 (spherical coordinates) showed that eφ(x) = rsinθ (r,θ,φ). In any event, en(x) varies with x, and one wonders just how en varies with x. For a small variation dx in the x-space coordinates, one has d(en)i = ∂j(en)i dxj ∂j ≡ ∂/∂xi or (den)i = (∂jen)i dxj Since en and (∂jen) are both vectors in ξ-space, and since the ek are known to form a complete basis in ξ-space, it must be possible to expand (∂jen) on the ek with some appropriate coefficients, call them Γkjn : (∂jen) = Γkjn ek Dotting this equation into ek gives Γkjn = ek (∂jen) = Rki(∂jRni) where we recall from Section 7 (s) that (ek)i = Rki and (en)i = Rni. These coefficients Γkjn(x) comprise a tensor-like field called the affine connection, so we shall regard the above line as the definition of Γkjn in the context of Picture C1 above. With more standard index names, the above becomes Γcab = ec (∂aeb) = Rci(∂aRbi) From Section 7 (q) we know that, for Picture C1, Rci = (∂xc/∂ξi) Rbi = (∂ξi/∂xb) (∂aRbi) = (∂2ξi/∂xa∂xb) = (∂bRai) Therefore one can write Γcab = Rci(∂aRbi) = (∂xc/∂ξi) (∂2ξi/∂xa∂xb) = which form appears in Weinberg p 100 (4.5.1). Notice again that (∂bRcn) = (∂cRbn) and that Γcbc is symmetric on the lower two indices. (b) Reverse-Tilt-of-R Derivative theorems The claims are : (∂aRdn) = – Ren Rdm (∂aRem) ∂a ≡ ∂/∂xa (∂aRdn) = – Ren Rdm (∂aRem) where on each line the R in one derivative has down-tilt indices and the other up-tilt indices. Either line can be obtained from the other by reflecting each R index pair in a horizontal line through the indices. Proof: These theorems are a simple consequence of the fact that RS = 1 which in standard notation is written δcb = RcαRbα (one of the orthogonality rules). So, 0 = ∂a(δde) = ∂a(RdmRem) = Rdm (∂aRem) + Rem (∂aRdm) (*) Apply Σe Ren to both sides of (*) to get 0 = Ren Rdm (∂aRem) + (Ren Rem) (∂aRdm) = Ren Rdm (∂aRem) + δnm (∂aRdm) = Ren Rdm (∂aRem) + (∂aRdn) => (∂aRdn) = – Ren Rdm (∂aRem) QED 1 Alternatively, apply Σd Rdn to both sides of (*) to get 0 = (Rdn Rdm) (∂aRem) + Rdn Rem (∂aRdm) = δnm (∂aRem) + Rdn Rem (∂aRdm) = (∂aRen) + Rdn Rem (∂aRdm) => (∂aRen) = – Rdn Rem (∂aRdm) now swap d and e: => (∂aRdn) = – Ren Rdm (∂aRem) QED 2 (c) Upper-index Γ theorem The claim is : gcb Γ dab + gdb Γcab = – (∂agcd) // sum on b where the two Γ objects have the same lower indices, but the upper index gets shuffled. Context is: Proof: First, the Γ objects can be replaced by their definitions. Γcab ≡ Rcn(∂aRbn) // from section (a) // ∂a = ∂/∂ξa Γdab ≡ Rdn(∂aRbn) // c→d The LHS of the claimed theorem may then be written LHSacd = gcb Rdn(∂aRbn) + gdb Rcn(∂aRbn) The metric tensors can be replaced by gcb = Rci Rbj Gij = Rci RbiGii gdb = Rdi Rbj Gij = Rdi RbiGii so that LHSacd = Rci Rbi Rdn(∂aRbn) Gii + Rdi Rbi Rcn(∂aRbn) Gii Now process the right side of the claimed theorem, – RHSacd = (∂agcd) = (∂a[Rci RdiGii]) = (∂a[Rci Rdi]) Gii = Rci(∂aRdi) Gii + Rdi(∂aRci) Gii Since the RHS and LHS involve R matrices of reverse tilts, the first "Reverse-Tilt-of-R Derivative theorem" of section (b) is recruited, (∂aRdn) = – Ren Rdm (∂aRem) (∂aRdi) = – Rei Rdm (∂aRem) // n→ i (∂aRci) = – Rei Rcm (∂aRem) // d → c Installing these last two lines into the RHS gives RHSacd = Rci Rei Rdm (∂aRem) Gii + Rdi Rei Rcm (∂aRem) Gii Now rename summation indices m→ n and e→ b RHSacd = Rci Rbi Rdn (∂aRbn) Gii + Rdi Rbi Rcn (∂aRbn) Gii and visual inspection shows that this is the same as LHSacd computed above, QED. (d) Theorem: Γdab = (1/2) gdc [ ∂agbc + ∂bgca – ∂cgab] The theorem claims that Γdab may be expressed entirely in terms of the metric tensor, Γdab = (1/2) gdc [ ∂agbc + ∂bgca – ∂cgab] Recall in our definition above, Γdab = Rdk(∂aRbk), that Γ was given in terms of R matrices. The following corollary ( derived at the end of this section ) concerns contraction of the upper Γ index with a lower one: Γaan = (1/2) gad ∂ngad = (1/2)(1/g)∂ng = (1/) ∂n() Proof: The quasi-Cartesian metric tensor G (see Section 1) is Gij = Gij = δij Gii were Gii = ±1 independently for each i We know that gab = RaeRbfGef = RaeRbe Gee gab = RaeRbfGef = RaeRbe Gee From the first of these lines, gdc = RdiRci Gii . The first line below is computed from the second line in the pair above, then the next two lines below are obtained by doing forward cyclic permutations of the first line : ∂cgab = [Rbe (∂cRae) + Rae(∂cRbe) ]Gee ∂agbc = [Rce (∂aRbe) + Rbe(∂aRce) ]Gee ∂bgca = [Rae (∂bRce) + Rce(∂bRae) ]Gee . The last four lines can be inserted into the Right Hand Side of our desired theorem to obtain (RHS)dab = (1/2) gdc { ∂agbc + ∂bgca – ∂cgab} = (1/2) RdiRci Gii Gee * [Rce (∂aRbe) + Rbe(∂aRce) + Rae (∂bRce) + Rce(∂bRae) – Rbe (∂cRae) – Rae(∂cRbe) ] 1 2 3 4 5 6 Due to the symmetry noted above in section (a), (∂iRje) = (∂jRie), terms 2 and 5 cancel as do terms 3 and 6, while terms 1 and 4 are equal. Therefore, (RHS)dab = (1/2) RdiRci Gii Gee * 2 Rce (∂aRbe) = Rdi (Rce Rci) Gii Gee (∂aRbe) = Rdi δei Gii Gee (∂aRbe) = Rde Gee Gee (∂aRbe) = Rde(∂aRbe) = Γdab QED The abovementioned corollary is this (g ≡ det(gab) ) Γaan = (1/2) gad( ∂agnd + ∂ngad – ∂dgan ) = (1/2) gad ∂ngad = (1/2) (1/g)∂ng = (1/) ∂n() where the first and third terms cancel due to symmetry gad∂agnd – gad∂dgan = gad∂agnd – gda∂agdn = gad∂agnd – gad∂agnd = 0 and the fact that (1/g) ∂ng = gab(∂ngab) is proved as follows: (1) gab = (g-1)ab = cof(gab)T/det(gab) = cof(gab)/g => cof(gab) = g gab (2) g = det(gab) = Σa gabcof(gab) => ∂g/∂gab = cof(gab) = g gab (3) ∂ng = ∂g/∂xn = (∂g/∂gab)( ∂gab/∂xn) = g gab (∂ngab) => (1/g) ∂ng = gab(∂ngab) The final form shown is just calculus : g-1/2 ∂n(g1/2) = g-1/2 (1/2) g-1/2 ∂n(g) = (1/2) (1/g) (∂ng) (e) Picture D1 Context. Our main context of interest is Picture A, . For the proof given in the next section below, it is useful to think of Picture A as the top part of this Picture D1, The relationship between the R's and S's are these R = R' R-1 = R' S => R' = RR S = R R'-1 = R S' The tangent and reciprocal base vectors in ξ-space associated with the transformations F and F' are these: (en)i = Rni (en)i = Rni x-space (e'n)i = R'ni (e'n)i = R'ni x'-space Now there are two affine connections, Γcab ≡ (∂xc/∂ξn) (∂2ξn/∂xa∂xb) = Rcn ∂a (∂ξn/∂xb) = Rcn(∂aRbn) ∂a = ∂/∂xa = [ec]i (∂a[eb]i) = ec (∂aeb) Γ'cab ≡ (∂x'c/∂ξn) (∂2ξn/∂x'a∂x'b) = R'cn ∂'a (∂ξn/∂x'b) = R'cn(∂'aR'bn) ∂'a = ∂/∂x'a = [e'c]i (∂'a[e'b]i) = e'c (∂'ae'b) (f) Relations between Γ and Γ ' The claimed relations are the following in the context of Picture A shown above, Γ'cab = Rcd Raα Rbβ Γdαβ + Rcα (∂'aRbα) // Weinberg (4.5.2) Γ'cab = Rcd Raα Rbβ Γdαβ – Rbβ(∂'aRcβ) Γ'cab = Rcd Raα Rbβ Γdαβ – Raα Rbβ (∂αRcβ) // Weinberg (4.5.8) If the second term were not present, the relation would state that Γdαβ transforms as a mixed rank-3 tensor in the usual manner (Section 7 (j)). Since the second term is present, Γdαβ is not a tensor. Proof of the first relation : We now make use of Picture D1 shown above. Start with the Γ' definition given above, Γ'cab ≡ R'cn(∂'aR'bn) = (RR)cn ∂'a(RR)bn = RcdRdn ∂'a(RbβRβn) // R' = RR = RcdRdnRbβ(∂'aRβn) + Rcd(RdnRβn)( ∂'aRbβ) = RcdRdnRbβ([Raα∂α]Rβn) + Rcd(δdβ)( ∂'aRbβ) // ∂'a = Raα∂α in first term only = RcdRaαRbβRdn(∂αRβn) + Rcβ (∂'aRbβ) = RcdRaαRbβ Γdαβ + Rcα (∂'aRbα) // Γdαβ ≡ Rdn(∂αRβn) QED Magically, all the R's have gone away. The second term in the above relation can be written a different manner as follows. Consider, 0 = ∂'a(δcb) = ∂'a(RcαRbα) = Rcα (∂'aRbα) + Rbα (∂'aRcα) => Rcα (∂'aRbα) = – Rbα(∂'aRcα) = – Rbβ(∂'aRcβ) = – Rbβ Raα(∂αRcβ) and this gives the other two relations stated above. (g) The covariant derivative of a vector is a rank-2 tensor The rest of this Appendix is in the Picture A context The "covariant derivative" of a vector is defined as Vβ;α ≡ [ ∂αVβ – Γ cαβ Vc] // in x-space V'b;a ≡ [∂'aV'b – Γ' cab V'c] // in x'-space The claim is that Vβ;α transforms as a covariant rank-2 tensor. One must then show that V'b;a = Raα Rbβ Vβ;α or [∂'aV'b – Γ' cab V'c] = Raα Rbβ [∂αVβ – Γ cαβ Vc] (*) Proof: The brute force method is to replace primed objects on the LHS with unprimed objects: (∂'aV'b) = (Raα∂α) (RbβVβ) = Raα Rbβ (∂αVβ) + Raα(∂α Rbβ)Vβ Γ'cab = Rcd Raα Rbβ Γdαβ + Rcα (∂'aRbα) // From section (f) above V'c = RceVe and ∂'a = Raα∂α The LHS of then becomes LHSb;a = [∂'aV'b – Γ' cab V'c] = Raα Rbβ (∂αVβ) + Raα(∂α Rbβ)Vβ – [Rcd Raα Rbβ Γdαβ + Rcα (∂'aRbα) ] RceVe = Raα Rbβ (∂αVβ) + Raα(∂α Rbβ)Vβ – (RceRcd) Raα Rbβ Γdαβ Ve – (RceRcα) (∂'aRbα) Ve] = Raα Rbβ (∂αVβ) + Raα(∂α Rbβ)Vβ – δedRaα Rbβ Γdαβ Ve – δeα (∂'aRbα) Ve] = Raα Rbβ (∂αVβ) + Raα(∂α Rbβ)Vβ – Raα Rbβ Γeαβ Ve – (∂'aRbe) Ve] // ∂'a = Raα∂α = Raα Rbβ (∂αVβ) + Raα(∂α Rbβ)Vβ – Raα Rbβ Γeαβ Ve – Raα(∂αRbe) Ve] The second and fourth terms cancel leaving = Raα Rbβ (∂αVβ) – Raα Rbβ Γeαβ Ve = Raα Rbβ [(∂αVβ) – Γcαβ Vc] = RHS QED Thus it has been shown that Vβ;α transforms as a covariant rank-2 tensor V'b;a = Raα Rbβ Vβ;α where Vβ;α ≡ [∂αVβ – Γ cαβ Vc] Summary. It immediately follows that all four of these transformations are valid (shown on the left) : V'b;a = Raα Rbβ Vβ;α Vβ;α ≡ [∂αVβ – Γ cαβ Vc] 1 V'b;a = Raα Rbβ Vβ;α Vβ;α = [∂αVβ + Γ βαc Vc] 2 V'b;a = Raα Rbβ Vβ;α Vβ;α = [∂αVβ – gαd Γ cdβ Vc] 3 V'b;a = Raα Rbβ Vβ;α Vβ;α = [∂αVβ + gαd Γ βdc Vc ] 4 We shall now derive the lower three expressions shown on the right. Start with Vβ;α ≡ [∂αVβ – Γ cαβ Vc] Apply Σβgbβ to both sides to get Vb;α = gbβVβ;α ≡ [gbβ(∂αVβ) – gbβΓ cαβ Vc] (*) The "Upper-Index Γ theorem" of section (c) states that gcb Γ dab + gdb Γcab = – (∂agcd) gcβ Γ bαβ + gbβ Γcαβ = – (∂αgcb) => gbβ Γcαβ = – gcβ Γ bαβ – (∂αgcb) // Γcβα = Γcαβ Installing this into (*) gives Vb;α = [gbβ(∂αVβ) + gcβ Γ bαβ Vc + (∂αgcb) Vc] = [gbc(∂αVc) + Γ bαβ Vβ + (∂αgcb) Vc] The first and third terms are recognized as the RHS of this evaluation (∂αVb) = ∂α(gbcVc) = (∂αgbc) Vc + gbc(∂αVc) and therefore Vb;α = [(∂αVb) + Γ bαβ Vβ ] QED 2 To this result apply Σαgaα to both sides to get Vb;a = gaα Vb;α = [gaα (∂αVb) + gaα Γ bαβ Vβ ] = [ (∂aVb) + gaα Γ bαβ Vβ ] QED 4 Starting over with Vβ;α ≡ [∂αVβ – Γ cαβ Vc] apply Σα gaα to both sides to get Vβ;a = gaαVβ;α ≡ [gaα(∂αVβ) – gaαΓ cαβ Vc] = [(∂aVβ) – gaαΓ cαβ Vc] QED 3 (h) Product rule and covariant derivative of a rank-n tensor Product Rule. Here are two examples of the product rule : (AaBb);n ≡ Aa;nBb + AaBb;n (AabcBde);n ≡ Aabc;n Bde + Aabc Bde;n More generally if A and B are arbitrary tensors each with an arbitrary set of up and down indices, then (A----B----);n ≡ A----;n B---- + A---- B----;n // Weinberg p 105 (4.6.14) where each tensor maintains all its indices and the ;n is simply distributed as shown. Once this fact is known, one can apply ;n to the product of more tensors, for example (A----B---- C----);n ≡ A----;n B---- C---- + A---- B----;n C---- + A---- B---- C----;n The implication here is that the object defined in this product rule transforms as a tensor in the obvious manner. For the cases shown above one would have (A'aB'b);n = Raa'Rbb' Rn'n(Aa'Bb');n' (A'abcB'de);n = Raa'Rbb'Rcc' Rdd'Ree' Rn'n (Aa'b'c'Bd'e');n' (A'----B'----);n = RRRR... RRRR... Rn'n (A-'-'-'-'B-'-'-'-');n' No proof of the product rule is provided here. The brute force method would be to write out the LHS and remove all primed expressions using known facts such as the relation between Γ' and Γ of section (f). Certainly there are more elegant methods. One might consider an induction proof if all else fails. Covariant derivative of a tensor of rank-n. Here are some examples of properly defined covariant derivatives of covariant tensors : Aab;α ≡ ∂α Aab – ΓnaαAnb – ΓnbαAan // Weinberg p 108 Babc;α ≡ ∂α Babc – ΓnaαBnbc – ΓnbαBanc – ΓncαBabn Notice that there is one Γ term for each index of the tensor being differentiated. The summation index n on the right "marches through" the indices of the tensor shown. The general case would be Babc..x;α ≡ ∂α Babc..x – ΓnaαBnbc..x – ΓnbαBanc..x – ........... – ΓnbαBacc..n (*) Again, the implication of this definition is that the resulting object transforms as a tensor: A'ab;α = Raa'Rbb'Rαα' Aa'b';α' B'abc;α = Raa'Rbb' Rcc'Rαα' Ba'b'c';α' B'abc..x;α = R R R R...... Rαα' Ba'b'c'..x';α' (**) The proof that (**) is indeed satisfied by (*) is left to the energetic reader. Appendix G: Expansion of (v) and div(T) in general curvilinear coordinates Although the polyadic notation is regarded as archaic by some writers (eg, Wolfram), it is well embedded into the literature of continuum mechanics, a field awash in rank-2 tensors. In this literature one sometimes sees (A)ij ≡ ∂jAi where the indices are the reverse of the normal dyadic definition of Appendix E because this makes certain equations look simpler. For example, in continuum mechanics one encounters the so-called convective or material derivative of an arbitrary vector field A(x,t) in the Eulerian or spatial "view" of the motion of a blob of continuous matter ( eg, Lai (3.4.3) and (3.4.8) ), DAi/Dt = ∂tAi + Ai v = ∂tAi + (∂jAi) vj = ∂tAi + (A)ij vj = ∂tAi + [(A) v]i => DA/Dt = ∂tA + (A) v Here v(x,t) is the velocity field of the moving matter blob. The notation DAi/Dt is just a historical notation for the total derivative dtAi = dAi(x,t)/dt. The matrix A is not a differential operator since the derivative does not act on the vector standing to the right of A, but one is still often interested in expressing A in curvilinear coordinates. We now carry out this task as an illustration of the various notations presented in Appendix E. We replace the generic vector A by generic vector v and this v has nothing to do with the v shown in the example above. As usual for the curvilinear coordinates application, x-space is taken to be Cartesian, and x'-space is that of the curvilinear coordinates x'. The object div(T) where T is a matrix will be described in section (f). Consider then the dyadic (v) ( a dyadic with indices reversed from the usual dyadic sense) (v)ij ≡ ∂jvi As shown in Appendix E, this dyadic can be expanded on the Cartesian x-space basis as (v) = Σcd(v)dc uducT = Σcd(∂cvd) uducT . The general plan for expressing (v) in curvilinear coordinates is this: (1) to cause uducT to be replaced by edecT, where en are the tangent basis vectors of the transformation from Cartesian x-space coordinates to curvilinear x'-space coordinates under the non-linear transformation x' = F(x). As shown in the various Sections of this document on differential operators, the en vectors in x-space (or their unit vector versions n) are the appropriate basis vectors for expressing any tensor (or tensor-like object) where x'-space coefficients are desired. That is to say, as discussed in Appendix E, A = ΣijA'ij eiejT (2) to write out (∂cvd) in terms of x'-space coordinates and objects (like v'i) The second part of this plan is quite simple so we do it will be done first. (a) Expressing ∂cvd in terms of x'-space objects One can express ∂cvd in terms of x'-space coordinates and objects in exactly the Christoffel manner of Section 7 (v). Using the covariant vector rule Va = RbaV'b shown in Section 7 (p) one easily gets ∂cvd = Σij (Ric∂'i)( Rjdv'j) = Σij Ric [(∂'iRjd) v'j + Rjd (∂'iv'j)] (*) Comment: As noted in Section 7 (v), for a non-linear transformation F the object ∂cvd is not a rank-2 tensor. If it were a rank 2 tensor, as in the case of linear F, it would be very easy to express ∂cvd in terms of x'-space objects and coordinates, according to the expansion of Appendix E, A = Σij A'ij ei ejT A'ij = RiaRjb Aab // A is a rank-2 tensor (v) = Σij (v)'ij ei ejT (v)'ij = RiaRjb(v)ab = RiaRjb∂bva = Rjb∂b Riava = (∂'jv'i) so that (v) = Σij (∂'jv'i) ei ejT and we are all done. When F is non-linear, as is the case for curvilinear coordinate transformations, the extra work shown below or its equivalent is required. (b) Expressing uducT in terms of eiejT In Section 3 (b) it was shown that e'n ≡ Ren = Σi Rin ei which in standard notation reads e'n = Σi Rin ei where (e'n)i = δni (*) This can quickly be verified by dotting both sides with em (the dual basis is a complete basis) e'n em = Σi Rin ei em LHS = e'n em = gab (e'n)a(em)b = gab δna(em)b = gnb(em)b = gnb(em)b = (em)n = Rmn RHS = Σi Rin ei em = Σi Rin δim = Rmn For present purposes, (*) can be re-expressed as, ud = Σe Red ee where (ud)e = δde Then ucT = Σf Rfc efT and so uducT = Σef Red Rfc ee efT and our section (b) task is completed. (c) Combining the two steps It has now been shown in section (a) that ∂cvd = Σij Ric [(∂'iRjd) v'j + Rjd (∂'iv'j)] and in section (b) that uducT = Σef Red Rfc ee efT Therefore, carefully doing one step at a time, (v) = Σcd(v)dc uducT = Σcd(∂cvd) uducT = Σcd { Σij Ric [(∂'iRjd) v'j + Rjd (∂'iv'j)] } { Σef Red Rfc ee efT } = Σcd Σef { Σij Ric [(∂'iRjd) v'j + Rjd (∂'iv'j)] } { Red Rfc ee efT } = Σcd Σef { Σij Rfc Ric [Red (∂'iRjd) v'j + Red Rjd (∂'iv'j)] } ee efT = Σef Σij ( ΣcRfc Ric) [Σd Red (∂'iRjd) v'j + (Σd Red Rjd) (∂'iv'j)] ee efT Since the metric tensor is in fact a tensor, one knows that g'fi = RfcRid gad . But g = 1 so this becomes g'fi = RfcRic and so one has ( ΣcRfc Ric) = g'fi (Σd Red Rjd) = g'ej which can be inserted into the above expression to give, = Σef Σij g'fi [Σd Red (∂'iRjd) v'j + g'ej (∂'iv'j)] ee efT Change ij to ab = Σef Σab g'fa [Σd Red (∂'aRbd) v'b + g'eb (∂'av'b)] ee efT and then change ef to ij = Σij { Σab g'ja [Σd Rid (∂'aRbd) v'b + g'ib (∂'av'b)] } ei ejT . Therefore, it has been shown that (v) = Σij Qij ei ejT = Σij [(v)(e)]ij ei ejT where Qij ≡ Σab g'ja [Σd Rid (∂'aRbd) v'b + g'ib (∂'av'b)] = [(v)(e)]ij Here the notation (v)(e) indicates, as suggested in Appendix E, that this matrix is associated with the ei basis which in turn is associated with the curvilinear coordinates x' . Since g' and R are functions of the curvilinear coordinates x', our program of expressing (v) entirely in terms of x'-space coordinates and objects is now complete. Notice that x' is an arbitrary not necessarily orthogonal coordinate system. (d) Alternate forms of the (v) expansion In terms of covariant components of vector v', it was shown above that Qij = Σab g'ja [Σd Rid (∂'aRbd) v'b + g'ib (∂'av'b)] (v) = Σij Qij ei ejT If one wants to see contravariant components of vector v', replace v'b = Σc g'bcv'c to get Qij = Σab g'ja [Σd Rid (∂'aRbd) Σc g'bcv'c + g'ib ∂'a(Σc g'bcv'c)] . Or, if one wants to see the unit-vector expansion components of vector v', replace v'c = h'c-1 v'c to get Qij = Σab g'ja [Σd Rid (∂'aRbd) Σc g'bc h'c-1 v'c + g'ib ∂'a(Σc g'bc h'c-1 v'c)] . As a reminder, these v'n arise as follows: v = Σn v'n en = Σn v'n (h'nn) = Σn (h'n v'n) n = Σn v'n n => v'n ≡ h'n v'n . As discussed in Section 14 Example 1, the components v'n are convenient since they all have the same dimensions. Moreover, when a specific curvilinear system is selected, one can dispense with the unpleasant font used in v'n and just write v'n = vx'(n). For example, in spherical coordinates r,θ,φ : v'1 = vr v'2 = vθ v'3 = vφ v = Σn v'n n = vr + vθ + vφ . Once one is talking v'n and n, one usually wants to see the (v) expansion in terms of i jT. Since ei ejT = h'ih'j i jT our expansion can be rephrased in this manner: (v) = Σij Qij ei ejT = Σij ( Qij h'ih'j) i jT = Σij Pij i jT with Pij = h'ih'jQij = [(v)(e^)]ij where Qij can be any of the three forms shown above. (e) Special case of orthogonal coordinates In this case g'nm = h'n2δn,m and g'nm = h'n-2δn,m . We consider only the last form for Qij above. Since the algebra is tedious, here are all the steps, where the next item to be processed is shown in red. Qij = Σab g'ja [Σd Rid (∂'aRbd) Σc g'bc h'c-1 v'c + g'ib ∂'a(Σc g'bc h'c-1 v'c)] = Σab h'j-2δj,a [Σd Rid (∂'aRbd) Σc g'bc h'c-1 v'c + g'ib ∂'a(Σc g'bc h'c-1 v'c)] = Σb h'j-2 [Σd Rid (∂'jRbd) Σc g'bc h'c-1 v'c + g'ib ∂'j(Σc g'bc h'c-1 v'c)] = h'j-2 [Σd Rid Σb (∂'jRbd) Σc g'bc h'c-1 v'c + Σb g'ib ∂'j Σc g'bc h'c-1 v'c)] = h'j-2 [Σd Rid Σb (∂'jRbd) Σc g'bc h'c-1 v'c + Σb h'i-2δi,b ∂'j(Σc g'bc h'c-1 v'c)] = h'j-2 [Σd Rid Σb (∂'jRbd) Σc g'bc h'c-1 v'c + h'i-2 ∂'j(Σc g'ic h'c-1 v'c)] = h'j-2 [Σd Rid Σb (∂'jRbd) Σc h'c2δb,c h'c-1 v'c + h'i-2 ∂'j(Σc h'i2δi,c h'c-1 v'c)] = h'j-2 [Σd Rid Σb (∂'jRbd) h'b v'b + h'i-2 ∂'j(h'iv'i)] = h'j-2 h'i-2 [h'i2Σd Rid Σb (∂'jRbd) h'b v'b + ∂'j(h'iv'i)] = h'j-2 h'i-2 [h'i2Σd Rid Σb (∂'jRbd) h'b v'b + (∂'jh'i)v'i + h'i(∂'jv'i)] so that then Pij = h'j-1 h'i-1 [h'i2Σd Rid Σb (∂'jRbd) h'b v'b + (∂'jh'i)v'i + h'i(∂'jv'i)] This object Pij can be computed in Maple by the following code, Computation of (grad v) in spherical coordinates This file is also the template for any other cases. > restart; > with(linalg): > N := 3; > xp[1] := r; > xp[2] := theta; > xp[3] := phi; > assume(r>0,theta>0,theta<Pi); > x[1] := r*sin(theta)*cos(phi); > x[2] := r*sin(theta)*sin(phi); > x[3] := r*cos(theta); > Vp := vector( [v[xp[1]],v[xp[2]],v[xp[3]] ] ); > S_ := (i,j) -> diff(x[i],xp[j]); > S := matrix(N,N,S_); R is not used anywhere, but we compute it anyway. > R := simplify(inverse(S)); > gcov := simplify(evalm( transpose(S) &* S)); > for n from 1 to N do hp[n] := simplify(sqrt(gcov[n,n])) od; > T1_ := (i,j) -> (hp[i])^2 * sum( sum(R[i,d]*Diff(R[b,d],xp[j])*hp[b]*Vp[b],d=1..N),b=1..N) ; > T2_ := (i,j) -> Vp[i] *diff(hp[i],xp[j]) ; > T3_ := (i,j) ->hp[i] * Diff(Vp[i],xp[j]) ; > P_ := (i,j) -> (1/hp[i])*(1/hp[j])*(value(T1_(i,j)) + (T2_(i,j)) + T3_(i,j)); > P :=matrix(N,N): > for n from 1 to N do > for m from 1 to N do > P[n,m] := expand(simplify(P_(n,m))); > od; > od; > evalm(P); which is easily modified for other orthogonal curvilinear systems. Here are some sample results: (v) = Σij Pij i jT Pij = [(v)(e^)]ij Pij in polar coordinates (where 1,2 = r,θ) : // agrees with Lai (2.23.23) Pij in cylindrical coordinates (where 1,2,3 = r,θ,z) : // agrees with Lai (2.34.5) The polar coordinates results are seen to be upper left 2x2 piece of the cylindrical results. Pij in spherical coordinates (where 1,2,3 = r,θ,φ) : // agrees with Lai (2.35.25) (f) Expansion of div(T) in general curvilinear coordinates The vector object A ≡ divT is defined in Cartesian coordinates by Ai = ∂jTij with implied sum on j. Regarding Ri as the ith row of the matrix Tij, this says that Ai = ∂j(Ri)j = Ri, so the components of divT are just the divergences of the rows of T. This divT object arises for example when Newton's 2nd Law F = ma is applied to a particle of continuous matter, in which case this law is known as Cauchy's Equation of Motion (Lai (4.7.4)) , divT + ρB = ρa where ρ is the mass density of the particle and B the action-at-a-distance force (body force) per unit mass. We wish to express divT in terms of curvilinear coordinates and objects. As usual, x-space is taken as a Cartesian space and x'-space as that of the curvilinear coordinates. The vector A = divT can be expanded according to Appendix E as A = Σi Ai ui = Σi (∂jTij) ui One can follow the same two-step plan as used above for (v) : ui = Σe Rei ee (∂kTij) = (Rak∂'a)(RbiRcjT'bc) = [Rak (∂'aRbi)RcjT'bc + Rak Rbi(∂'aRcj)T'bc + Rak RbiRcj (∂'aT'bc) ] Comment: Note that (∂kTij) would be a tensor if the first two terms above vanished, but of course they don't vanish for general underlying transformation F to curvilinear coordinates so (∂kTij) is not a rank-3 tensor, and Σj(∂jTij) is not a tensorial vector. So Ai = [divT]i = ∂jTij is just a "vector-like" object. Under normal rotations, however, divT would be a vector, as suggested by divT + ρB = ρa. Inserting both the above expressions into the expansion of A and using Rac Rbc = g'ab in various places, one obtains, A = Σe Qe ee where Qe = [Rei (∂'aRbi) g'ac T'bc + Raj g'eb(∂'aRcj)T'bc +g'ebg'ac (∂'aT'bc) ] For an orthogonal x' coordinate system Qe becomes Qe = [Rei (∂'aRbi) h'a-2 T'ba + Raj h'e-2 (∂'aRcj)T'ec + h'e-2h'a-2 (∂'aT'ea) ] Using ee = h'e e, the orthogonal result can be restated as A = Σe Pe e where Pe = h'e [Rei (∂'aRbi) h'a-2 T'ba + Raj h'e-2 (∂'aRcj)T'ec + h'e-2h'a-2 (∂'aT'ea) ] In the above, it has been assumed that Tij is a rank-2 tensor so that, in the notation of App. E, T = ΣijTij(uiuj) = ΣijT'ij(eiej) If one is interested in an expansion of T on the unit vectors ei = h'i i this becomes T = Σij[T'ijh'ih'j] (ij) = Σij[T(e^)]ij (ij) One then has [T(e^)]ij = h'ih'jT'ij => T'ij = [T(e^)]ij/ (h'ih'j) => T'ij = [T(e^)]ij (h'ih'j) where the last result follows from T'ij = g'ii'g'jj'T'ij = h'i2h'j2 T'ij. When this last result is inserted into Pe one gets Pe = [Rei (∂'aRbi) h'a-1 h'b h'e [T(e^)]ba sum on i, a, b + Rai (∂'aRbi) h'b [T(e^)]eb sum on i, a, b + h'e-1h'a-2 ∂'a(h'e h'a) [T(e^)]ea sum on a + h'a-1(∂'a[T(e^)]ea) sum on a The above object Pe can easily be computed by Maple code similar to that shown in the (v) case and here are some results: Pi ≡ [divT]i in cylindrical coordinates (where 1,2,3 = r,θ,z) : // agrees with Lai (2.34.8,9,10) For polar coordinates P1 and P2 are given by the first two lines above with the last terms set to 0. These polar results then agree with Lai (2.33.32,33). Pi ≡ [divT]i in spherical coordinates (where 1,2,3 = r,θ,φ) : The expressions above agree with Lai (2.35,33,34,35). References L. A. Ahlfors, Complex Analysis, 2nd Ed. ( McGraw-Hill, New York, 1966). G. Backus, Continuum Mechanics (Samizdat Press, Golden Colo., 1997). J.D. Bjorken and S.D. Drell, Relativistic Quantum Mechanics (McGraw-Hill, New York, 1964). E. B. Christoffel, "Ueber die Transformation der homogenen Differentialausdrücke zweiten Grades", Journal für die reine and angewandte Mathemarik, 70 (1869), 46–70, 241–245. This paper may be found in Christoffel's Collected Mathematical papers, Gesammelte Mathematische Abhandlungen, 2 vols. (Tuebner, Leipzig-Berlin, 1910), downloadable from Google books. R. Hermann, Ricci and Levi-Civita's Tensor Analysis Paper (English translation with comments) (Math Sci Press, Brookline, MA, 1975, perhaps on-line) First, with much praise, Hermann provides an English translation of the French Ricci & Levi-Civita paper referenced below, updating words, phrases and symbols to the current day. Second, he inserts perhaps 180 pages of inline italicized text which translates the ideas of the paper into mathematical frameworks not known or not connected to by the authors (eg, fiber bundles, direct product spaces, Killing vectors, moving frames, group theory, etc.). Hermann presents material that was more simply understood by later authors (eg, E. Cartan). Earlier in 1966 Hermann wrote a book Lie Groups for Physicists which contains the group theory chunk of this added material. One realizes that differential geometry is a very large field touching upon many areas of Mathematics (a house with many mansions). M. Lai, E. Krempl and D.Ruben, Introduction to Continuum Mechanics, 4th Ed. (Elsevier, Amsterdam, 2009). H. Margenau and G.M. Murphy, The Mathematics of Physics and Chemistry, 2nd Ed. (D. van Nostrand, London. 1956). A. Messiah, Quantum Mechanics (John Wiley, New York, 1958). Reference is made to page 878 (Vol II) of the North-Holland 1966 fifth printing paperback two-volume set, Chapter XX paragraph 2. P. Moon and D.E. Spencer, Field Theory Handbook, Including Coordinate Systems, Differential Equations and their Solutions (Springer-Verlag, Berlin, 1961). This is the place to go to find explicit expressions for differential operators in specific curvilinear coordinate systems (and very much more). P.M. Morse and H. Feshbach, Methods of Theoretical Physics ( McGraw-Hill, New York, 1953). M.M.G. Ricci, T. Levi-Civita, "Méthodes de calcul différentiel absolu et leurs applications", Mathematische Annalen (Springer) 54 (1–2): 125–201 (March 1900). This huge 77 page paper is sometimes referred to as "the bible of tensor analysis". Levi-Civita was a student of Ricci and they worked together on this paper and later elaborations. Their work was instrumental in Einstein's later discovery of general relativity. The title is "Methods of absolute differential calculus and their applications". Absolute differential calculus was the authors' phrase for what is now called Tensor Analysis/Calculus/Algebra. The word absolute referred to the idea of equations being covariant (see Section 7 (u) above). I. Stakgold, Boundary Value Problems of Mathematical Physics, Volumes 1 and 2 (Macmillan, London, 1967). J.J. Sylvester, "On the General Theory of Associated Algebraical Forms" (Cambridge and Dublin Math. Journal, VI, pp 289-293, 1851). This paper appears in H.F. Baker, Ed., The Collected Mathematical Papers of James Joseph Sylvester (Cambridge University Press, 1901). S. Weinberg, Gravitation and Cosmology: Principles and Applications of the General Theory of Relativity John Wiley & Sons, New York, 1972). E. B. Wilson (notes of J.W. Gibbs), Vector Analysis (Dover, New York, 1960)