Phil Lucht Math & Physics Archive
Home / Math and Physics Files / Math / Curvilinear Systems / Tensor Doc and Support / the previous tensor paper

tensors and curvilinear coordinates

DOCX · 410.4 KB
Open DOCX file

Phil's expository paper dated 11.16.09, with a note from 11.3.11 saying it contains errors, was superseded by a newer and longer tensor document, and is kept for historical purposes only. It builds tensor algebra from coordinate transformations, covers covariant and contravariant tensors, raising and lowering indices, and the metric tensor. It then derives base vectors, volume elements, and the divergence, gradient, Laplacian and curl in general curvilinear coordinates, with examples and appendices.

AI-written summary; may contain errors.

Extracted text (machine-read; may contain errors)
Tensors and Curvilinear Coordinates PhL 11.16.09 [ I reviewed this doc on 11.3.11 and found a few items of interest which got added to the new tensor doc. There are various errors below left in place, and the whole approach is somewhat incorrect, but I think it was a very good shot, if not ready for prime time. The newer doc is twice as long and much more coherent and I think is ready for prime time. I plan to retain this doc for historical purposes only, and would not recommend anyone read it. ] See Preface below. Preface 3 1. Transformation Tab (down-tilt form). 4 Example A. 2D Rotation. 4 Example B. Polar to Cartesian. 5 2. Contravariant Tensors 5 3. Transformation Tab (up-tilt form). 6 4. The orthogonality rules, matrix forms, and the inverse transformations 7 (a) The orthogonality rules. 7 (b) Matrix notation. 7 (c) The Inverse Transformations. 8 5. Covariant Tensors 9 6. Tensors and the idea of Raising and Lowering Indices 9 (a) Lowering matrix gab 9 (b) Raising matrix gab 9 (c) Matrix g as a tensor. 10 (d) Action of g' in x' space. 10 (e) g acting on itself . 10 (f) A new relation between Tab and Tab. 11 (g) The contraction tilt rule. 11 (h) Summary to this Point: 12 7. A candidate for g : the metric tensor 12 (a) x-space as more than just a vector space. 12 (b) Example 1: Rotations 13 (c) Example 2: Lorentz Transformations. 14 (d) Summary: x-space is now a real Hilbert Space . 14 8. Interlude 15 9. Application to curvilinear coordinates 16 (a) A formula for g'ab 16 (b) A summary of results including matrix notation. 17 (c) Notations used for metric tensor diagonal elements. 18 (d) Example C: 2D Cartesian to Polar coordinates 18 10. The tangent base vectors ei and their relation to the metric tensor. 20 Some definitions: Level surfaces, level curves and coordinate lines. 20 (a) The x-space tangent base vectors. 21 (b) The q-space tangent base vectors. 23 (c) The tangent base vector parallelepiped and its volume v. 23 (d) Carl Gustav Jacob Jacobi (1804 –1851). 25 (e) Expressing v in terms of g'ab: new symbol g'. 25 (f) Example: Spherical Coordinates 26 11. The reciprocal base vectors ei . 28 (a) The Reciprocal Base Vectors in N dimensions. 28 (b) The Reciprocal Base Vectors in 3D. 30 (c) A confusing notational issue concerning vectors. 31 (d) Vectors expanded on the base vectors in various ways. 33 (e) Verify that both definitions of the ei agree; the generalized cross product 34 12. The transformation under T of differential distance, area and volume elements. 36 (a) Distance. 36 (b) Area. 36 (c) Volume. (N=3 and general N) 37 (d) Summary of the results. 38 13. The Divergence in curvilinear coordinates. 38 (a) Definition of div A. 39 (b) Calculation of the surface integral. 39 (c) Solve for div A 40 (c) Various forms of div A. 41 (c) div A in orthogonal systems. 41 14. The Gradient in curvilinear coordinates. 42 (a) Definition of grad(f) 42 (b) Calculation of grad(f) 42 (c) Various forms of grad(f) 42 (d) The combination A grad(f) 43 15. The Laplacian in curvilinear coordinates. 44 16. The Curl in curvilinear coordinates. 44 (a) Definition of curl A. 44 (b) Computation of the circulation. 45 (c) Solving for curl A. 46 (d) various forms of curl A . 47 (e) curl A in orthogonal systems. 47 17. The Vector Laplacian in curvilinear coordinates. 48 18. Summary of the differential operators in curvilinear coordinates. 50 (a) for general systems 50 (b) for 3D orthogonal systems 51 (c) Continued Example: Spherical Coordinates. 51 Appendix A. Matrix notation. 53 Appendix B. Three questions of possible interest. 55 (a) Is Tab a mixed tensor? 55 (b) Can we raise and lower indices on Tab ? 55 (c) Are the objects δab and εabc.. tensors? 57 Appendix C. The Action of a Gradient on a Vector 57 (a) Preliminary comments: 57 (b) Gradient of a Vector: Elwin Bruno Christoffel (1829–1900). 57 (c) Some technical Facts. 58 Preface The subjects of curvilinear coordinates and tensor algebra are connected, the former being an application of the latter. The term "curvilinear coordinates" usually refers to one of the 11 different orthogonal 3D coordinate systems (the classical systems) in which the Helmholtz (and thus Laplace) equation can be "separated", a subject discussed at length in Morse & Feshbach 1953 (M&F). Two familiar examples are spherical and cylindrical coordinates. A less familiar example is ellipsoidal coordinates, in which the potential of a charged ellipsoidal conductor can easily be calculated, and from which all of the other 10 classical coordinate systems can be derived as degenerate cases. Our discussion below, however, shall include all possible coordinate systems, not just orthogonal ones. It is the subject of tensors and their behavior under transformations that is the main concern below. This subject is always confusing to a student and even to many old timers I suspect. The up and down indices always seem strange. Why must we have such things, one wonders? We don't need such things for Cartesian systems, but we do for any other system since the metric tensor is then not the identity. The entire tensor subject below is constructed from the notion of a vector and how it transforms under a change of coordinates. Group theory is only mentioned in passing in a few places, but in many applications a vector is associated with a certain representation of a group, like the rotation group ( j = 1) or the Lorentz group (1/2 1/2). Both groups have spinor representations (j = 1/2 for rotations, and 0 1/2 and 1/2 0, dotted and undotted, for Lorentz) and these spinors can be used as the basic objects of the algebra, in which case a whole similar universe can be constructed under the name spinor algebra, as compared with the tensor algebra described below. Apart from this comment, we avoid this subject of spinors. I have used various curvilinear systems for decades, but only in August 2008 (age 60) did I really sit down and try to understand the subject. Since I owned a dusty copy of Margenau and Murphy (M&M) "The Mathematics of Physics and Chemistry", 2nd Ed. 1956, and since that book contained this subject, that is where I learned the basics. M&M have separate discussions of tensor algebra (a tiny part of Chap 4) and curvilinear coordinates (all of Chap 5), but they don't really connect these subjects tightly, though something is said in the last section of Chap 5. I have tried to make that connection clearer. They do catalog a few basic facts about each of the 11 commonly used coordinate systems. The 15 chapters of M&M cover, in about 600 pages, a huge spectrum of classical applied mathematics. M&F have 2000 pages for their presentations. In my discussion below, I construct the world of tensor algebra using an arbitrary raising and lowering matrix, and then I treat the metric tensor as a useful candidate for this matrix. My fundamental notation for the transformations is Tab for contravariant vectors, and Tab for covariant ones. M&M write out the cross derivatives everywhere and it just seems better to have a more compact notation. I do not recall much from my brief study of general relativity, so I don't know exactly how the tensor subject ties in, but I know it does, and in a big way. I think one of the main points is that the metric tensor is indeed a function of the coordinates, and that makes the whole thing run. Mass affects the metric tensor, causing space to be a curvilinear coordinate system instead of a Cartesian one. The general flow of the presentation below is simple. First we talk tensors, then we get into the curvilinear coordinate systems as a "tensor application", and generate general formulas for operators like divergence. Subjects out of the main flow are mentioned in a few appendices at the end. This is my first-ever document to have an "outline" in Word that is not a jumbled mess, a comment the reader should ignore. 1. Transformation Tab (down-tilt form). A transformation from one set of independent coordinates xi to another independent set x' i is defined by a transformation matrix T made of partial derivatives. At its most fundamental level, this matrix tells us how differential coordinates transform: (repeated indices are summed unless otherwise noted) dx'a = dxb ≡ Tab dxb Tab = (*) [ In this document, we regard the xi and x' i and the matrix elements of T as being real numbers. ] The location of the vector and matrix indices (their up and down sense) will be addressed below. Since the first index on T is up and the second is down, we refer to this as T in down-tilt form. The leftmost equality on the line above is nothing more than the derivative chain rule for function x'a(x1, x2, ...), and we have then taken the derivative as the definition of Tab. No rocket science here. Under this transformation T, an object like dxb transforming as shown above is called a contravariant vector relative to T. All the notions of scalars, vectors, tensors, covariant and contravariant (to be discussed below), are specific to a transformation T. Cartesian coordinates xi and their differentials dxi (both contravariant vectors) are the everyday coordinates people use in physics. When T is a rotation in normal E3 space, it turns out that the up or down position of the indices on vectors or on T makes no difference, so people usually put the indices down so the index is not confused with an exponent. The number of xi coordinates is N. Most results presented below are for arbitrary N (perhaps N ≥ 2), but examples will be used which have N = 2, 3 or 4. Results valid only for N=3 will be so marked. Example A. 2D Rotation. Here is an immediate T example, the "active" z rotation of a 2D vector r = (x,y) = r' = Rz(θ) r dr' = Rz(θ) dr x' = cosθ x - sinθ y y' = sinθ x + cosθ y T11 = ∂x'1/∂x1 = ∂x'/∂x = cosθ T12 = ∂x'1/∂x2 = ∂x'/∂y = -sinθ T21 = ∂x'2/∂x1 = ∂y'/∂x = sinθ T22 = ∂x'2/∂x2 = ∂y'/∂y = cosθ Tab = This example illustrates several points. A transformation T can have one or more "parameters" like θ shown here. For each θ we have a different transformation. For each such transformation, dr transforms as a contravariant vector according to the definition given above. We might indicate this transformation by the notation T(θ). One might think of the set of rotations over the range of the parameter(s) to be a certain class of transformations and consider them as such. The matrix elements Tab are "constants" in our example, meaning they are not themselves functions of the coordinates (we show a counterexample next). For this reason, not only does dr transform as a vector, so does r. Given dr' = T dr we integrate both sides (since T is constant) to get r' = T r . Since T is independent of the variables, we note that r' = T r is a linear transformation, T(r1 + r2) = Tr1 + Tr2. In general, T will not be constant, and none of these facts will be true. Example B. Polar to Cartesian. Here at once is an example of a non-linear transformation, one that takes us from polar coordinates to 2D Cartesian coordinates: x = (x1,x2) = (r,θ) x' = (x'1,x'2) = (x',y') x' = rcosθ dx' = cosθ dr – rsinθ dθ y' = rsinθ dy' = sinθ dr + rcosθ dθ T11 = ∂x'1/∂x1 = ∂x'/∂r = cosθ T12 = ∂x'1/∂x2 = ∂x'/∂θ = -rsinθ T21 = ∂x'2/∂x1 = ∂y'/∂r = sinθ T22 = ∂x'2/∂x2 = ∂y'/∂θ = rcosθ Tab = dx' = T dx In our previous example, θ was a parameter, but here it is a coordinate. We see that coordinates are appearing inside the matrix T, so T is non-linear and cannot be "integrated" in the simple sense mentioned above. Thus we have dx' = T dx but x' ≠ T x ( for example x' ≠ cosθ r – rsinθ θ ) The first example above is a prototype for spatial rotations or Lorentz transformations. The second example is a prototype for the type of transformation encountered in the study of "curvilinear coordinates". It should be stressed that the transformation from x to x' can be arbitrarily complicated, apart perhaps from issues of confused topology or degenerate axes. One should not think of T as being "just rotations". 2. Contravariant Tensors A contravariant tensor with respect to transformation T is an object that transforms under T in a certain specific manner. Here are some examples (these are definitions of the object types named) [ We apologize for the fact that primes ' appear on objects like S' and also on indices like a'. These primes are completely unrelated. I find the index primes useful to see where things connect, and the object prime is embedded in the whole presentation, so I did not want to give up either use. Greek indices are harder to read in my usual font and Word display factor (122%) so I avoided them too. ] S' = S contravariant tensor of rank 0 ( "scalar") V'a = Taa' Va' contravariant tensor of rank 1 ("vector") X'ab= Taa' Tbb' Xa'b' contravariant tensor of rank 2 fix up subscripts etc U'abc= Taa' Tbb' Tcc'Ua'b'c' contravariant tensor of rank 3 etc. The object with the prime, like S' or V', is the object as it appears in the space x', whereas objects like S and V are objects as they appear in space x. These spaces are related by dx'a = Tab dxb and by the integral of this equation, which is some x' = f(x). This is the passive view of the transformation. In the active view, something like V' is how V appears in the original space but after it has been transformed. For example, if we have v' = Rv , then we can think of v' as being the vector v as seen in x' space, or as a vector v' in x-space which has been rotated some amount and no longer aligns with v. We shall have much to say on this subject later. An example of a contravariant vector would be a coordinate xμ or momentum pμ with respect to a Lorentz Transformation, and one would say x'μ = Λμνxν . As already noted above, for 3-vectors in E3 space, it turns out the up and down index sense makes no difference, so we usually see x'i = Rijxj. There are of course various examples of rank 0 tensors, such as constants or the dot product of two vectors, and of rank 2 tensors in E3 such as the inertia, polarization, or stress tensors. The objects with contravariant indices are the usual objects one encounters in physics, but there are exceptions such as the gradient operator noted below. Apart from these exceptions, it is useful to think of the contravariant objects as being the "real" objects, and the other objects we shall construct below as being dependent on one's choice of the matrix gab to be discussed below. 3. Transformation Tab (up-tilt form). So far we have dealt with only the down-tilt form of the matrix T, where we refer to the up/down location of indices on T. To review, above we had: dx'a = dxb ≡ Tab dxb Tab = (*) down-tilt We noted that something that transforms like dxb is a contravariant vector relative to T. One can separately define an "up-tilt" version of T as follows: (note that the primed coordinate is now on the bottom of the derivative). dx'a = dxb ≡ Tab dxb Tab = (**) up-tilt As a guide, notice that in either tilt form, the first index of T matches that of the primed coordinate. That is to say, the first index of T is associated with the x' space, while the second with the x space. Something that transforms like dxa is called a covariant vector relative to T and this thing is generally not the same as dxa, more on this below. Notice that both indices appearing in the derivatives defining T are contravariant (up) since these are the normal coordinates one uses. An example of an normal object that transforms as a covariant vector is the gradient of a scalar, or just the gradient operator ∂a ≡ ∂/∂xa. We have ∂'a ≡ ∂/∂x'a = = = ∂b so ∂'a = Tab ∂b which is again just the differentiation chain rule in action. 4. The orthogonality rules, matrix forms, and the inverse transformations (a) The orthogonality rules. It seems pretty clear that the following is true: = δac = 1ac (ie, matrix elements of the identity matrix) this being just the definition of a partial derivative in a set of independent coordinates. If we throw in the chain rule on the left this becomes = δac = 1ac (ie, matrix elements of the identity matrix) or Tba Tbc = 1ac ≡ δac " orthogonality rule #1 for T" We could do the chain rule the other way to say = δac = 1ac (ie, matrix elements of the identity matrix) or Tab Tcb = 1ac ≡ δac " orthogonality rule #2 for T" So the "orthogonality rules" involves the product of reverse-tilted T matrices which are summed either on both first indices or on both second indices. To summarize : sum on firsts sum on seconds δac = Tba Tbc = Tbc Tba = Tab Tcb = Tcb Tab orthogonality rules (b) Matrix notation. The idea here is to define matrices with all lower indices so we don't get confused about indices being up or down. Perhaps this is superfluous, but it helps the author a lot. ( More detail is given in Appendix A. ) If we indicate the transpose of a matrix M by MT ( this T has nothing to do with the transformation T), then we can define some "up" and "down" conventional matrices and their transposes as follows (Tup)ab ≡ Tab = (Tup)Tba (Tdn)ab ≡ Tab = (Tdn)Tba Then the orthogonality rules just derived can be written in this matrix form: 1 = (Tup)T Tdn = (Tdn)T Tup = Tdn(Tup)T = Tup (Tdn)T Recall that det(ab) = det(a)det(b) and also det(aT) = det(a). Thus, we know that 1 = det(Tup)det(Tdn) Therefore, if we knew that Tup was invertible, we would also know that Tdn was invertible. In this case we could conclude that Tdn = (Tup-1)T = (TupT)-1 and Tup = (Tdn-1)T = (TdnT)-1 which says that the either matrix is the inverse of the transpose of the other matrix. This is then a definite statement of the impression we get that the two T's appear to be inverses in some sense, from their definitions: Tab ≡ Tab ≡ We can now go back to "tilted index notation" and write Tdn = Tup-1,T => Tab = (T-1,T)ab = (T-1)ba Tup = Tdn-1,T => Tab = (T-1,T)ab = (T-1)ba (c) The Inverse Transformations. We have just shown that (T-1)ab = Tba and (T-1)ab = Tba Now here is the contravariant transformation and its inverse dx'a = Tab dxb => dxa = (T-1)ab dx'b To verify the right equation, apply Tca from the left to both sides and get dx'a = dx'a. But then we use our inverse forms just stated to get dxa = dx'b Tba . We can process the covariant equation in a similar fashion, and here then are the results: transformation inverse transformation contravariant dx'a = Tab dxb => dxa = (T-1)ab dx'b = dx'b Tba covariant dx'a = Tab dxb => dxa = (T-1)ab dx'b = dx'b Tba 5. Covariant Tensors Just as we extended the rule dx'a = Tab dxb to define contravariant tensors: S' = S contravariant tensor of rank 0 V'a = Taa' Va' contravariant tensor of rank 1 X'ab= Taa' Tbb' Xa'b' contravariant tensor of rank 2 U'abc= Taa' Tbb' Tcc'Ua'b'c' contravariant tensor of rank 3 so also we can extend the rule dx'a = Tab dxb to define covariant tensors: S' = S covariant rank 0 ( same as contravariant rank 0) V'a = Taa'Va' covariant rank 1 X'ab= Taa' Tbb' Xa'b' covariant rank 2 U'abc= Taa' Tbb' Tcc'Ua'b'c' covariant rank 3 We refer to an upper index as a contravariant index, and a lower one as a covariant index. Finally, we can define mixed tensors where some indices are contravariant and others covariant. In this case, each index gets transformed with its appropriately tilted T matrix U'abc= Taa' Tbb'Tcc' Ua'b'c' mixed rank 3, two contravariant, one covariant Again, we are just saying that if something is a mixed tensor of this type, then this is how it transforms, by definition. 6. Tensors and the idea of Raising and Lowering Indices (a) Lowering matrix gab Let gab be ANY symmetric invertible matrix. We will then use it to convert a contravariant vector into a covariant one (working in the x space here) Va ≡ gabVb // lower an index on a contravariant vector [ But if Vb is a contravariant vector, how do you know Va will be a covariant one? This is not going to be true for gab being ANY symmetric matrix. My approach of thinking of the metric tensor as just a candidate g is this disproven. (b) Raising matrix gab We would like to find some corresponding other matrix which we will call gab which undoes what gab does. We would like to say for example Vb = gbaVa so gba raises a covariant index, making it contravariant. This then implies that Vb = gbaVa = gba gab'Vb' => gba gab' = δbb' = 1bb' Therefore, it must be true that the matrix gab is the inverse of the matrix gab. gab = (g-1)ab // therefore gab is also symmetric [ gup = gdn-1] (c) Matrix g as a tensor. We can cause our symmetric invertible matrix gab to be a rank 2 tensor if we define g' in the x' space in this manner g' ab = Tac Tbd gcd contravariant tensor of rank 2 g' ab = Tac Tbd gcd covariant tensor of rank 2 Note that g' is in general going to be different from g. (d) Action of g' in x' space. The matrix (now tensor) g' in x' space raises and lowers indices on x'-space tensors in the same way that g raises and lowers indices on x-space tensors. For example, we will have V' a ≡ g'abV 'b (*) V' b = g' baV'a g 'ab = (g' -1)ab Note that g' is defined as in (c) and then it is used in x' space to generate newly defined objects like V'a. One can easily verify the first two equations above (and other like it) from facts previously stated. For example, the RHS of equation (*) is this: g'abV 'b = (Tac Tbd gcd) (TbeVe) = Tac(TbdTbe) gcdVe = Tac δde gcdVe = Tac gcdVd = TacVc = V'a (e) g acting on itself . Since g is a tensor, we can use g to raise and lower indices on itself as we do on any other tensor. For example gabgbc = gac But we know that gab is the inverse of gab, so it must be true that gabgbc = gac = δa,c One usually writes the Kronecker delta in the notation δac since it keeps track of index locations. That is, we might use δac = δac = δac as needed. So one more time: gabgbc = gac = δac (f) A new relation between Tab and Tab. Now let's look at our definitions from Sections 1 and 3 applied to arbitrary contravariant and covariant vectors V'a = Tab Vb Tab = (*) V'a = Tab Vb Tab = (**) We now process the first equation (*) in this manner V'a = Tab Vb (g' abV'b) = Tab (gbcVc) g'ea(g' abV'b) = g'eaTab (gbcVc) V'e = (g'eaTab gbc) Vc But on comparison with our second equation (**) transcribed as V'e = Tec Vc we find that (g'eaTab gbc) = Tec which gives us a relationship between the up-tilt and down-tilt form of the T matrix. Admittedly, it is a little complicated, since it involves both g and g'. The relations between Tab and Tab obtained in Section 4 above are simpler and don't involve g or g'. Those relations, which we optionally wrote in matrix form, are really statements of the orthogonality rules. See Appendix B for discussion of the indices on T. (g) The contraction tilt rule. Here we just want to do some fiddling with the g matrices to develop a "tilt rule". We will do this entirely in our x space with objects which exist in x space. Consider this quantity, where we have one "contracted index" b, AabBbc We can write this as (notice that we quietly make use of the fact that gab is symmetric) AabBbc = (gbb'Aab')( gbb" Bb"c) = (gbb' gbb") Aab' Bb"c = δb'b" Aab' Bb"c = Aab Bbc or AabBbc = Aab Bbc This shows the "tilt rule" for a contracted index. The rule is that you can take any summed pair of indices and change their "tilt" without changing the value of the expression of interest. Here is another example: XaYa = XaYa The contraction of a pair of indices "on a tilt" reduces the rank of a tensor by 2. Whereas XaYb is a rank two tensor ( this particular form is called an outer product), the contraction XaYa is a rank zero tensor which is called a scalar. It is a scalar because it is invariant under the transformation T: X'aY'a = (TabXb)(TacYc) = (TabTac) XbYc = δbc XbYc = XbYb So one benefit of having a matrix gab is that it allows us to construct tensors of reduced rank by contracting larger tensors (which requires one of the two indices of a pair to be lower). In particular, it allows us to construct scalars. A common notation for our scalar is this: X.Y = XbYb = XbYb = gabXaYb = gabXaYb It should be emphasized that everything being said here "floats" on the selection of gab. If you change gab, then you change the scalar X.Y, since it is equal to gabXbYb (presumably the two contravariant vectors here are "physical objects" which are independent of gab). (h) Summary to this Point: We have constructed a whole world of tensor algebra based on an arbitrary symmetric and invertible matrix gab and a transformation T. Each choice of gab gives a whole different tensor algebra world "beyond" the contravariant tensors that we start with. In the next section, we make a particular choice of gab that is applicable to a class of x-spaces known as metric spaces. Footnote: The reader might be wondering about the tensor nature of the object Tab and whether or not it makes sense to raise and lower its indices. This subject is addressed in Appendix B. In what follows, we shall not do any such raising or lowering, and so Appendix B can be ignored. 7. A candidate for g : the metric tensor (a) x-space as more than just a vector space. A linear (aka vector) space is one in which there are rules for adding and scaling vectors, but there is not necessarily any rule for describing the length of a vector. One can add such a rule called a norm, and say that ||x|| is then the length of vector x. One then has a normed linear space. But we still don't have any way to talk about the distance between two vectors. We can then add a metric function d such that d(x,y) is the "distance" between two vectors. Once choice of metric, known as the "natural metric induced by the norm", is to define d(x,y) = ||x-y||. Now we have a metric space in which we can talk about both the length of vectors and the distance between vectors. In our discussion above, we have not talked about such matters, we just assumed some generic vector space we called x-space (and another one x'-space) and we formed a tensor algebra world using some tensor gab and its inverse gab to raise and lower indices. We now want to consider our x-space to be a normed linear space in which the norm (squared) of a vector v is defined as follows: || v ||2 ≡ Gabvavb (*) ( norm = "length" of a vector) where Gab is some "norm matrix" which characterizes our normed linear space. For Cartesian space, we would have G = 1. If G were not symmetric, by the way, we could write it as G = Gsym + Gantisym and the second matrix would make no contribution to || v ||2, so we might as well assume G is symmetric. Our "norm matrix" determines the distance between two vectors in this way d(u,v) = || u- v || = Gab (u-v)a(u-v)b ( metric = "distance" between two vectors) so then we might refer to our symmetric Gab as a "the metric matrix". In particular, the distance between two points in x-space is given by d(x1,x2) = || x1- x2 || = Gab (x1-x2)a(x1-x2)b We can now consider Gab as a "candidate" for the index-lowering tensor gab we used in our tensor algebra discussion. We would then find that the length of a vector is invariant under transformation T, which is to say, the length is a tensorial scalar, || v ||2 ≡ Gabvavb = gabvavb = vbvb = v.v = a tensorial scalar (*) = v'.v' = g'abv'av'b = G'abv'av'b = || v' ||2 where now the x'-space is also a metric space whose metric matrix is G'ab = g'ab. The implication here is that no matter how horribly complicated our transformation T might be, the norm or "length" of a vector will be invariant under the transformation. In particular, the "distance" between two points in x-space will be the same as the "distance" between the corresponding two points in x'-space, where distance is defined as above. This is a very desirable situation to have be true, and that is why we are going to approve our candidate for office. By identifying our metric matrix in x-space with the index lowering tensor, and similarly in x' space, we may now change the name of g = G to be "the metric tensor", which it is commonly called. We could apply (*) to a position vector x or its differential dx, in which case we get || dx || 2 = gabdxadxb => || dx' || 2 = || dx || 2 ≡ (ds)2 where is called "the invariant distance". So under any transformation T, (ds)2 does not change. (b) Example 1: Rotations Consider regular old E3 space with metric tensor g = diag(1,1,1) so g-1 = g. We then have || x || 2 = gabxaxb = xaxa = x2 + y2 + z2  = r2 = || x' || 2 = x'2 + y'2 + z'2  = r'2 For T = rotations R, we have x'a = Tabxb = Rabxb. We can then show that g' = g, g' ab = Rac Rbd gcd = Rac Rbc = RacRbc = RacR-1cb = δab = gab since rotations are real orthogonal transformations RT = R-1. If we considered T = weird skew transformation on E3, then we will still have || x || 2= || x' || 2, but we will NOT have g' = g. Instead, we will have g' ab = Tac Tbd gcd = Tac Tbc . This is precisely the situation encountered when we define a curvilinear coordinate system. Irrelevant note added: For a rotation by θ about axis we can write Tki = [R()]ki = [exp(-iJ)]ki where (Ji)jk = -i ijk. where Ji are one form of the three generators of the vector representation of the Rotation Group. (c) Example 2: Lorentz Transformations. Consider 4D Minkowski space with vectors xμ = (ct, x) and with g = diag(1,-1,-1,-1) and hence g-1 = g. With this metric tensor we have || x || 2 = gabxaxb = (c2t2- x x ) For T = Lorentz transformations, we will find that g' = g by an argument similar to that given above for rotations, so we will have (ds)2 = (c2dt2- dx dx ) = (c2dt'2- dx' dx' ) = (ds')2 In special relativity, this "invariant distance" ds is called "the proper time", often written dτ . Irrelevant note added: The general Lorentz Transformation can be written Tki = [exp(± i σμν Mμν)] ki where (Mμν)kp = gμkgνp - gνkgμp where the matrices Mμν are one form of the six generators of the vector representation of the Lorentz Group. They are constructed from the elements of the metric tensor g = diag(1,-1,-1,-1). The σij are rotation parameters and σk0 are boost parameters. σμν and Mμν are both antisymmetric. (d) Summary: x-space is now a real Hilbert Space . If our x-space is not simply a bland vector space but is a metric space having a metric matrix Gab (the vector space is "endowed" with a metric), then, if we identify our raising tensor gab in x-space with Gab, we shall find that the x'-space is also a metric space with metric matrix G'ab = g'ab . Moreover, we shall find that the "length" (squared) of any vector, as defined by gabvavb, will be the same in both spaces. Without the identification of g = G, we would have v.v ≠ |v|2. Naturally we want to have v.v = |v|2 , so we could have started off at the beginning defining gab = Gab, as many authors do. We just wanted to show that the whole tensor world, useless as it might be without the norm = scalar connection, could be constructed from any symmetric invertible matrix gab. It might be added that u.v = gabuavb can be regarded as an "inner product" (u,v) = u.v of two vectors (in x-space). Since our x-space with g = G is then using || u ||2 = (u,u) and d(u,v) = || u - v || (the natural norm and the natural metric), our x-space is in fact a real Hilbert space, in which all vectors are real, so we have (u,v) = u.v = v.u = (v,u). In this entire document, x and x' are regarded as real coordinates. We could generalize to complex coordinates by doubling the number of real coordinates as in z = x + iy, or perhaps in some other way, but we shall not digress into that subject. 8. Interlude In sections prior to Section 7 above, we considered an arbitrary transformation T from one space to another which we characterized by the way the differential dxα coordinates transformed under T dx'a = dxb ≡ Tab dxb We considered examples in which T was a constant, and in which T depended on the coordinates. In the first example T was linear, in the second example, not. We then regarded the above transformation rule as applying to any contravariant vector V, V' a = Tab Vb and we extended this idea to define the transformation of contravariant tensors of arbitrary rank. We then considered a different form of our transformation dx'a = dxb ≡ Tab dxb which acted on what we called covariant vectors, and we identified the gradient operator as such a vector. We then took an arbitrary symmetric invertible matrix gab and its inverse gab and we used this to construct new tensor objects by raising and lowering indices on existing objects. For example, we could take the physical covariant gradient vector ∂a and create a new contravariant object ∂a = gab∂b from it. And we could take a physical vector va and create a new object va = gabvb from it. In this way, we constructed a whole tensor algebra world, starting from our known physical tensors. We noticed that if we contracted a covariant vector with a contravariant one, the resulting object was a tensorial scalar that did not change under the transformation T. Examples would be ∂ava or dxadxa. Along the way, we found that in the transformed space (the x' space) the raising/lowering tensor was g'ab obtained from gab by the same rule that transforms any rank 2 tensor. Each choice of gab generated a whole different tensor algebra world. We then decided to consider a special type of vector space known as a metric space and we then decided to select, as the raising/lowering tensor g of our tensor algebra, the metric tensor G of the metric space. So after doing this, the entire tensor algebra world is "carried" on the metric tensor G, which we then just called g. Change the metric tensor and you change the tensor algebra world -- all the non-physical objects change. When we made the association g = G, we found the distance (metric) between two neighboring points in x-space ( which we called ||dx||) could be identified with a tensorial scalar object of our tensor algebra. We decided this might be a "useful" concept, and discussed two examples where we managed to, in this way, connect the tensorial scalar property with a Cartesian distance in E3 and then with the proper time of special relativity in Minkowski space. We are now going to delve into a perhaps more mundane topic, that of curvilinear coordinates obtained by some transformation T from Cartesian coordinates. We shall normally only deal with contravariant vectors and in general we won't care about all the tensor algebra objects. However, we shall be interested in the metric tensor g' in the transformed space, given that g = 1 in the Cartesian space. We shall find that g'ab = Tac Tbc and is thus something easily computed for any given transformation T. Each different transformation T (there are an infinite number of them) defines a new curvilinear space. We shall find that the object g'ab tells us the relationship between certain interesting vectors in the x' space, which will be of great interest to us. Then we shall go on to discover the form of objects like gradient, divergence and curl as they appear in the x' space. In the case of rotations and Lorentz transformations, we were interested in transformations that were all of a certain class (perhaps all the transformations of a class formed a group), and each transformation could be described by some parameters like T(a,b) for a Lorentz transformation, where a and b are certain rotation and boost parameters. In dealing with curvilinear coordinates, we are more interested in a fixed transformation T that does not have any such parameters. Along the same lines, in the rotation and relativity examples, the T transformations do not depend on the coordinates and are thus linear. In contrast, our curvilinear T transformations will always depend on the coordinates and will always be non-linear. We might just mention while "interluding" that, in the case of a set of linear transformations that form a group, such as the Rotation Group SO(3) or the Lorentz Group SO(3,1), the vectors can be associated with a certain "vector representation" of the group. In the rotation case, a vector va is associated with the j = 1 representation of SO(3), and in the Lorentz group case the vector aμ is associated with the 1/2 1/2 representation of SO(3,1). In this context, the various tensors of the tensor algebra can be associated with direct product representations of these groups, and a bit more life is thereby pumped into the tensor algebra world. For example, these tensors can be "reduced" to a sum of "irreducible" representations. An example is the tensor Qij which is a 1 1 direct product representation of the rotation group, so it can be decomposed somehow into pieces which belong to the 2,1 and 0 representations. 9. Application to curvilinear coordinates (a) A formula for g'ab In this application of the above ideas, we always start in a Cartesian space xa with g = diag(1,1,1...) . We then move to some other x' space with coordinates x'a ≡ qa where I choose the letter q just because it is used by Margenau and Murphy. The x'a ≡ qa coordinates are the "curvilinear" coordinates. We shall use the terms q-space and x'-space interchangebly. In q-space, our metric tensor is some g' , and in x-space it is g = 1. We can compute g' from its assumed tensor transformation property g' ab = Tac Tbd gcd = Tac Tbd δcd = Tac Tbc which involves what I like to call the "x'-space indices of T" which are the first indices. Recall that Tab ≡ first index is the x' space index So in this application where g = 1 we have a formula for the metric tensor in q-space g' ab = Tac Tbc // sum on second index which, by the way, is manifestly symmetric. We conjecture now that g'ab = Tac Tbc We can verify this be remembering that g'ab and g' ab are inverses of each other, so let's check δab = g'ac g'cb = (Tad Tcd )( Tce Tbe) = Tad (Tcd Tce) Tbe = Tad δde Tbe = Tae Tbe = δab where we have used the orthogonality rule twice. Just for fun, we can check on our previously obtained relation between the two tilts of T, which was (g'eaTab gbc) = Tec (*) We can evaluate the left and side as follows (g'eaTab gbc) = (g'ea)Tab (gbc) = (TedTad)Tab (δbc) = (TedTad)Tac = Ted(TadTac) = Ted δdc = Tec and so (*) is then verified. So our tilts in this case are related by Tec = g'ea Tac [valid for g = 1, something wrong above? ] (b) A summary of results including matrix notation. Here then we summarize our results: g'ab = Tac Tbc or g'up = Tdn(Tdn)T g'ab = Tac Tbc or g'dn = Tup(Tup)T Tec = g'ea Tac (ds)2 = g'abdqadqb // all in q space (ds)2 = dxadxa // as ds appears in x space It is useful to rewrite the first two items above in the matrix notation introduced in Section 4 above, g'ab = Tac Tbc ↔ g'up = Tdn(Tdn)T = (Tup-1)T (Tup-1) g'ab = Tac Tbc ↔ g'dn = Tup(Tup)T = (Tdn-1)T(Tdn-1) In general, our curvilinear transformation equations will always be in contravariant coordinates, and from these we can therefore compute a Tdn = Tab type matrix. We have no way to directly compute the Tup matrix elements unless someone hands us the transformation equations in covariant coordinates which is not going to happen. So our USEFUL equations above are these g'up = Tdn(Tdn)T g'dn = (Tdn-1)T(Tdn-1) g'dn = (g'up)-1 where the last equation is just a reminder. So you can compute either g' matrix using one of the first two lines, then get the other one by inversion on the third line. We are thinking computer algebra programs here like Maple or MATLAB which are very happy to transpose and invert matrices. More vector/matrix/dyadic notation can be used for some of the results above. It is all the same thing, but when using Maple some notations help more than others. (ds)2 = g'abdqadqb = (dqT )a (g'dn)ab (dq)b = ( x x) = dqT g'dn dq = dq g'dn dq It is very tempting to just remove the prime from g', but I find doing that to be confusing in the larger picture of things, so we shall keep the prime right through the bloody end! ( q-space is x'-space) (c) Notations used for metric tensor diagonal elements. In an orthogonal curvilinear coordinate system, the metric tensor g'ab is diagonal. The diagonal elements are of course then just g'aa . As we see later, when differential operators like divergence are expressed in the x' = q coordinates, the numbers are what appear. Margenau and Murphy use Qa2 = g'aa so that Qa = . A common notation is ha = and ha is called a "scale factor" for reasons we shall see later. (d) Example C: 2D Cartesian to Polar coordinates The example will be the prototype for all our curvilinear work. We consider here the reverse of Example B of Section 1 above. That is to say, we want to look now at the transformation from 2D Cartesian to Polar coordinates. But first, recall example B ( Polar to Cartesian) which we summarize here from above: x = (x1,x2) = (r,θ) x' = (x'1,x'2) = (x',y') x' = rcosθ dx' = cosθ dr – rsinθ dθ y' = rsinθ dy' = sinθ dr + rcosθ dθ Tab = dx' = T dx If we swap primed with unprimed, we get our Example C situation of Cartesian x to Polar x' = q : x' = (x'1,x'2) = (q1,q2) = (r,θ) x = (x1,x2) = (x,y) x = rcosθ dx = cosθ dr – rsinθ dθ y = rsinθ dy = sinθ dr + rcosθ dθ The correct T for example C, which we shall call TC, is the inverse of T for example B, which is just called T. That is, (TC)ab = (T-1)ab dx' = TC dx = T-1 dx The g'ab metric tensor is then g'dn = TC,up (TC,up)T = (TC,dn-1)T (TC,dn-1) = (Tdn)T Tdn = [ transpose of ] [] = = That is to say, we label our matrix in the order 1,2 = r,θ and we have g'ab = = If we want to know g'ab it is of course just the inverse of g'ab so we have g'ab = = Note: Even in an orthogonal system, generally g'ab ≠ g'ab. From the first result we can construct our "invariant distance" (ds')2 = g'ab dx'a dx'b = 1 (dr)2 + r2 (dθ)2 Back in x-space, this is (ds)2 = 1(dx)2 + 1 (dy)2 Since this quantity is a scalar relative to TC (or T), it is the same in both spaces, so we conclude that 1 (dr)2 + r2 (dθ)2 = 1(dx)2 + 1 (dy)2 which is the distance between two neighboring points. We can confirm this result directly from the equations x = rcosθ dx = cosθ dr – rsinθ dθ y = rsinθ dy = sinθ dr + rcosθ dθ 1(dx)2 + 1 (dy)2 = (cosθ dr – rsinθ dθ)2 + (sinθ dr + rcosθ dθ)2 = (dr)2 + r2(dθ)2 10. The tangent base vectors ei and their relation to the metric tensor. Our pictures are drawn for 3D, but everything applies to general N dimensions. Some definitions: Level surfaces, level curves, level sets, and coordinate lines. I added these definitions on 10.27.10 because I was confusing level curves with coordinate lines. This short section contains information not presented earlier in this doc which may not be familiar to the reader. As an example, consider a family of confocal ellipsoids aligned with the axes. Each such ellipsoid has a label ρ and an equation x2/ρ2+ y2/(ρ2-h2)+ z2/(ρ2-k2) = 1. If we were to solve the cubic equation in ρ2 this represents to find ρ(x,y,z), then ρ(x,y,z) = ρ1 describes the ellipsoidal surface whose label is ρ1. This surface is an example of what is called a level surface. I would also call this ellipsoid a surface of constant parameter, in that case that parameter being ρ. We take a particular value of ρ, and let the other variables vary full range to get our surface. In general, any f(x,y,z) = const defines a level surface in 3D. If we take a z = z1 slice through our family of confocal ellipsoids, we get a family of confocal ellipses with these equations: x2/ρ2+ y2/(ρ2-h2) = 1 - z12/(ρ2-k2) ≡ a2. Then x2/(aρ)2+ y2/[(aρ)2-(ah)2] = 1 and we can just write this as x2/ρ'2+ y2/(ρ'2-h'2) = 1. Each ellipse has focus h' = ah, and label ρ'. These ellipses for varying ρ' are examples of a set of level curves. We could solve our ellipse equation for ρ'(x,y) and then each of these ellipses is described by ρ'(x,y) = const. In general, f(x,y) = const defines a set of level curves as const varies. Another example of level curves occurs in Ahlfors's chapter on complex mappings u+iv = F(x+iy). If we set u = const (and vary v full range), we get a set of level curves in the w plane, one for each const. We could write these as Re[F(x+iy)] = const, so in this case we have f(x,y) = Re[F(x+iy)]. Another orthogonal set of level curves is of course provided by v = constant, in which case f(x,y) = Im[F(x+iy)]. More generally, these things f(x,y....) = constant are called level sets. Now, what happens in a 3D curvilinear coordinate system if we start at some point and then vary q1 but we hold the other two coordinates q2 and q3 constant? In x,y,x space we move along a certain curved line in 3D space. This line in general does not lie in a plane (though usually it does). According to our definition above it would be incorrect to refer to this line as a "level curve" because it is not some f(x,y) = const. One can think of q2 = const as a level surface, and q3 = const is then some other level surface. The locus of both being constant is an intersection of these two surfaces, and that is our funny curved line. The correct name for such a line is a coordinate line. In almost all coordinate systems, the coordinate lines are in fact planar. For example, in spherical coordinates if you vary φ, holding r and θ fixed, you get a circle in the z=constant plane. I think in the ellipsoidal system, however, the coordinate lines are non-planar. I tacked some code onto the end of the file "ECSURF ellipsoidal coords surfaces.mws" to make it plot a a piece of a coordinate line. As I rotate the plot in Maple, I cannot make the curve appear as a line segment. The best I could do was this which tells us that indeed, this piece of coordinate line is non-planar. This is true despite the fact that the ellipsoidal coordinate system is orthogonal. One other fact: in an orthogonal system, the geometry near a Point is always "locally Cartesian", and this means that the q1 coordinate line leaves this Point in a direction normal to the surface of constant q1. As we shall see below, that surface at that Point contains two "tangent base vectors" e2 and e3 which are perpendicular to each other and to e1 which points along the q1 coordinate line at the Point. That is why we are claiming that in an orthogonal system, the tangent base vector ei is perp to the surface of constant qi. An example would be er in spherical coordinates being normal to a surface of constant r (sphere). (a) The x-space tangent base vectors. As noted above, we will always start in Cartesian coordinates xa and end up in some curvilinear coordinates x'a ≡ qa, where we choose the letter q because M&M do. Our opening gambit is the following picture, with explanation following. It needs to be emphasized that everything in this picture -- the axes, the various curves and vectors like e3 -- are being drawn in Cartesian space; think of it as real physical space. Footnote: An equation like dx = T-1dq below is an abbreviation for either of these two equations: dxk = (T-1)kjdqj dxk = (T-1)kjdqj So, the transformation from x'-space (= q-space) to x-space (Cartesian) says dx = T-1dq where T-1 is a square matrix. We can integrate this to get x = f(q) where f is some complicated vector function of q. If we start with some value of q in q-space and then move along one of the axes in q-space (the above picture does not show q space), we get some corresponding movement of x in x space. For the q1 axis we could write this as x + Δx = f (q + Δq1 1). As we vary Δq1 from perhaps -1 to +1, x + Δx traces out a curved line in x-space which passes through point x. This is the q1 "coordinate line" shown in the above drawing. We would like to find a vector tangent to this curve at the point x. Let's first write x + dx = f (q + dq ) = f(q) + (dq q) f(q) or dx = (dq q) f(q) which we can then apply to our situation dx = (dq1 1 q) f(q) = (∂q1f(q)) dq1 so we can define our tangent vector as follows e1 ≡ (∂q1f(q)) This "tangent base vector" e1 appears in the drawing. We then have (e1)j = ∂xj/∂q1 = ∂xj/∂x'1 = T1j and of course we can do this as well for Δq moving along the other axes, so we have (ei)j = ∂xj/∂qi = ∂xj/∂x'i = Tij and we find that, in fact, the tangent base vectors ei are the rows of the matrix Tij = Tup. The ei are in general not unit vectors, so they have no "hats". Also, the three ei vectors are in general not orthogonal. In an "orthogonal curvilinear coordinate system" (like spherical coordinates, see below), they will be orthogonal, but for now we assume the general case where they are not. Notice that the ei vectors are tangent to their curves at a point of interest x = f(q). As we move to some other point of interest, the vectors ei are likely to swing about wildly to different positions. For N > 3 dimensions, we have to imagine a picture with more dimensions to it, but the tangent base vectors are defined exactly as shown above, ei ≡ ∂x/∂qi. The metric tensor we know transforms as any other rank 2 tensor, g' ab = Tac Tbd gcd , and since our x-space is Cartesian, gcd = δc,d and we have g' ij = Tik Tjk. Therefore, we know that ei ej = (ei)k (ej)k = Tik Tjk = g' ij so we could write out the entire q-space metric tensor as a set of dot products of tangent base vectors. If this metric tensor is diagonal, then the tangent base vectors are orthogonal (and vice versa). (b) The q-space tangent base vectors. Although we shall not use these vectors, it should be clear that we can reverse the sense of the above discussion. We draw this corresponding picture: [To save effort for the graphic artist, we have shown the curves and vectors below graphically unchanged from their appearance above, but the reader will understand that the curves in q-space and their tangent vectors are of course completely different.] Everything drawn here is in q-space. The function g is f-1 of the previous picture. We end up with some q-space base vectors e'i which will be these: (e'i)j = ∂qj/∂xi = ∂x'j/∂xi = Tji = (T-1)ij So, no surprise, the reverse vectors e'i are the rows of the inverse T matrix, and are the columns of the matrix Tji = Tdn. Therefore we could write (just a notation fiddle exercise) (e'i)j = (e'up)ij = Tji = (Tdn)ji = (TdnT)ij When we refer to tangent base vectors below, we shall always mean the x-space ones ei . Notational warning Consider an x-space tangent base vector en . This is a genuine issue vector in Cartesian space with contravariant components (en)i. If we state how this vector should transform to be a vector, we would write, (en')i = Tij (en)j, this being the rule for how contravariant vectors transform. The vector obtained in this fashion is not the same as our x'-space tangent base vector which we have called by the exact same name en' . To see that this is the case, consider (en')i ≡ Tij (en)j = Tij Tnj = δin (en')iq-space thing = Tin ≠ δin The first line shows that the x-space tangent base vectors en , when transformed to q-space, are just unit vectors along the q-space axes, where we could call them perhaps n . The second line objects are NOT these unit vectors. Luckily we are never going to do anything with the q-space tangent base vectors. (c) The tangent base vector parallelepiped and its volume v. In our x-space picture above we showed the tangent base vectors as tangent to the coordinate lines at a point of interest. Here we draw those same base tangent vectors ei at the same point of interest, and we also draw the parallelepiped which they span. Let's call the volume of this parallelepiped v. If we were to map a q-space volume dq1dq2dq3 near point q into a volume in x-space near x, we would get in x-space a scaled copy of the above parallelepiped with edges dqi ei. The volume of this scaled parallelepiped would then be dVx-space = dq1dq2dq3 v. If one is doing an integration and changing variables, one can then make the replacement dx dy dz = dVx-space = v dq1dq2dq3 For example, if the qi are spherical coordinates, we have dx dy dz = v drdθdφ = r2sinθ drdθdφ so it must be that v = r2sinθ. We shall verify this in the section (f) below. Solid geometry tells us (the reader may have to ponder this) how to determine v from the base vectors ( we write this lots of ways): ±v = [e1, e2, e3] ≡ e1 e2 x e3 = + cyclic = - anticyclic = - swap [We put the ± because e1 e2 x e3 might be negative depending on the relative orientation of the three vectors. The symbol "v" is always meant to be a positive number. It seems easier to use ± with this understanding, rather than add absolute value marks everywhere. Warning: this was a "repair" done on 11.14.10 after I saw e1 e2 x e3 coming out negative in Maple for oblate coordinates. I have not really gone through everything below in this doc to make this repair. ] The cross product e1 e2 x e3 can also be written in this manner e1 e2 x e3 = (e1)i [e2 x e3]i = (e1)i εijk (e2)j (e3)k = εijk (e1)i(e2)j (e3)k which we recognize as the determinant of a matrix whose three rows are e1, e2 and e3. But we know (see section(a) above) this matrix is Tij = Tup. Since det(M) = det(MT), v must also be the |determinant| of the matrix whose columns are e1, e2 and e3 . So we summarize ways to write the volume for our 3D case : ±v = det(Tij) = det(Tup) = det [ (e1) , (e2), (e3) ] = εijk (e1)i(e2)j (e3)k It will not surprise the reader to know that this line of equalities extends to N dimensions in this way ±v = det(Tij) = det(Tup) = det [ (e1) , (e2), (e3), ..... ] = εijk.... (e1)i(e2)j (e3)k ..... (*) where dots mean continue out to N indices or objects. The case N=2 is not hard to verify: ±v = εij (e1)i(e2)j = (e1)1(e2)2 - (e1)2(e2)1 If e1 is placed at the origin pointing to the right along , and e2 at some angle θ, this becomes ±v = (e1)1(e2)2 = |e1| |e2| sinθ = ± the area of a parallelogram One could imagine some kind of induction proof of (*) for doubtful readers. It is hard to imagine the result could be anything other than as shown, given that it works for N = 2 and N = 3. The exact same argument for the transformation of a differential volume under transformation T, presented above in 3D, applies in N dimensions, and we have [dx1dx2dx3...... ] = v [ dq1dq2dq3..... ] ±v = det(Tab) = det () = det J(x,x') where J(x,x') is called the Jacobian Matrix, and is just another name for our = Tab , used in the context of volume element transformations. Accepting the above formula for volume in N dimensions, we have just proved that the volume transformation in going from one coordinate system to another is given by det J(x,x') defined as shown above. Since 1 = (Tup)T Tdn from orthogonality, we know that det(Tab) = 1/det(Tab), so we can write ±1/v = det(Tab) = det() = det J(x',x) which gives us another way to compute v. (d) Carl Gustav Jacob Jacobi (1804 –1851). German, Jewish, Berlin PhD 1825, then Konigsberg. Elucidated the whole world of elliptic integrals and functions, such as F(x,k) and sn(x;k), which occur even in simple problems like the 2D pendulum. Wiki claims he promoted Legendre's ∂ symbol for partial derivatives (used throughout this document) and made it a standard. Among many other contributions, he saw the significance of the object J(x',x) which now bears his name. The Jacobi Identity is another familiar item, a rule for non-commuting operators [x,[y,z]] + [z,[x,y]] + [y,[z,x]] = 0 which finds use in quantum mechanical operators and matrices, and more generally with Lie group generators. (e) Expressing v in terms of g'ab: new symbol g'. This seems the right place to insert this important result. Earlier we showed that (ei ej) = g'ij , so we can write det(g') = εijk... g'1ig'2jg'3k... = εijk... (e1 ei) (e2 ej) (e3 ek) .... = εijk.... (e1)a (ei)a (e2)b (ej)b (e3)c (ek)c ..... We know that (ei)a = Tia = (Tup)ia. Let S ≡ Tup so we have det(g') = εijk... S1a Sia S2b Sjb S3c Skc .... On the other hand, we know that ±v = det(Tab) = det(Tup) = det(S) so v2 = det(S) det(S) = det(S) det(ST) = det(SST) But we know that (SST)ij = SiaSTaj = Sia Sja so v2 = det(SST) = εijk.... (SST)1i(SST)2j(SST)3k.... = εijk.... S1a Sia S2b Sjb S3c Sjc... But this is exactly the expression given for det(g') above. Thus we have shown that v = In our discussion of differential operators below, we shall define the following symbol g' ≡ det(g') => v = = We could of course define g ≡ det(g) = det(1) = 1, so we never use such a g. It is true that the symbol g' is now "overloaded", to borrow a software term, since it is sometimes a matrix and sometimes a number, but the meaning is always clear from the context. We only do this to make results below more compact. Notice that det(g') is written out in terms of lower index gab elements, so it is really det(g'dn). Since det(M-1) = 1/det(M), we know that det(g'up) = 1/ det(g'dn) . So g' = det(g') ≡ det(g'dn) det(g'up) = 1/ det(g'dn) = 1/g' (f) Example: Spherical Coordinates We set (x1, x2, x3) = (x,y,z) (q1, q2, q3) = (θ,φ,r) // note that r is last here! x = r sinθ cosφ y = r sinθ sinφ z = r cosθ (ei)j = ∂xj/∂qi = ∂xj/∂x'i = Tij We shall enter our data into Maple and have it compute the three tangent base vectors (as rows of the Tij = Tup matrix), and then have it compute v = e3 e1 x e2 . The reader unfamiliar with Maple should still be able to read this code fairly easily: which shows that v = r2sinθ as claimed earlier. Dredging up some external data on spherical coordinates, we find that ( ) = = Rz(φ) Ry(θ) where the columns of the matrix shown here are the unit vectors shown. Comparing these columns to the ei vectors from Maple (which are the rows of Tup) , we see that e1 = r e2 = r sinθ e3 = e1 x e2 = r2sinθ x = r2sinθ v = e3 e1 x e2 = r2sinθ So in this case, we find that our tangent base vectors are not normalized (except for e3) and are mutually orthogonal. The e3 = = r/r is dimensionless, which explains why our volume has dimension L2. Of course the combination r2sinθdrdθdφ has dimensions L3. Now let's have Maple compute the metric tensor and its determinant, using the fact that ei ej = (ei)k (ej)k = Tik Tjk = g' ij so that g' ij = Tik Tjk = (Tup)(TupT)ij We have already computed Tup in the above code, so we continue to get which says that g' = diag( r2, r2sin2θ, 1) = diag(Qθ2, Qφ2, Qr2) in terms of the "scale factors": Qθ = r Qφ = rsinθ Qr = 1 // often called hθ,hφ, hr We don't need Maple to see that det(g') = r4sin2θ which is indeed v2. 11. The reciprocal base vectors ei . (a) The Reciprocal Base Vectors in N dimensions. The reciprocal base vectors ei are "reciprocal to" the tangent base vectors ei defined above. These reciprocal vectors are defined on the first line below, then on the second line we compare this to our definition of the tangent base vectors (ei)j ≡ Tij reciprocal base vectors ei is the ith row of Tdn (ei)j ≡ Tij tangent base vectors ei is the ith row of Tup Sometimes the set of vectors ei is called "the dual basis" to ei. In Appendix B we talk about raising and lowering indices on the T matrix, something we have not done much. The first index is associated with x' space and moves up and down under the action of g', while the second index is associated with x space and it moves up and down with g. But in our curvilinear coordinates discussion, we have gab = δa,b, so we can move the second index up and down for free. Then we have, for example, Tij= Tij and Tij = Tij. The implication of this fact for our two kinds of base vectors is this: (ei)j= (ei)j (ei)j= (ei)j So all these vectors are the same whether in covariant or contravariant components. Thus, we can write, for example e1. e2 ≡ (e1)k (e2)k = (e1)k (e2)k = e1 e2 so we can always use the non-tensorial Cartesian dot product notation for such dot products. [ Perhaps this is obvious since we said these vectors are all vectors in Cartesian space. ] Let's now define matrices eup and edn according to (edn)ik = (ei)k = Tik = (Tdn)ik up and down refer to the TILT of the indices (eup)ik = (ei)k = Tik = (Tup)ik so we have merely given our T matrices new names: edn = Tdn eup = Tup so edn eupT = 1 eupT = edn-1 etc [ Recall that the q-space tangent base vectors were (e'i)j = Tji so we do not have e'i = ei since we know that (ei)j ≡ Tij. Rather, we see that the ei are the rows of the Tdn matrix while the q-space e'j are its columns. Also, needless to say, the reciprocal vectors are different from the vectors obtained by transforming the regular tangent base vectors to q-space, which transformed vectors we showed above are just axis-aligned unit vectors in q-space. So these en reciprocal vectors are something completely new. ] When expressed in terms of the two kinds of base vectors, T orthogonality says this, ei ej = (ei)k (ej)k = Tik Tjk = δij edneupT = 1 ei ej = (ei)k (ej)k = Tik Tjk = δij eupednT = 1 and, similarly, we may recast our metric tensor results in this vector notation ei ej = (ei)k (ej)k = Tik Tjk = g'ij eupeupT = g'up (g'up)ab ≡ g'ab ei ej = (ei)k (ej)k = Tik Tjk = g'ij ednednT = g'dn (g'dn)ab ≡ g'ab Also, based on the comment above about raising and lowering T indices, we have (ei)j = Tij = g'ikTkj = g'ikTkj = g'ik (ek)j (ei)j = Tij = g'ikTkj = g'ikTkj = g'ik (ek)j so ei = g' ik ek and ei = g' ik ek which can also be shown just from the matrix equations involving eup, edn, g'up and g'dn. More notational clarification. The last equations above suggest that ek is a some kind of covariant object whose corresponding contravariant object is ek. Let's make a little comparison between two equations: (v')i = g' ik (v')k // raising an index on vector v', something we can do in q-space = x'-space ei = g' ik ek // the first equation above In the first line g' is rasing a vector index from lower to upper, but in the second line g is raising the label on a vector from a lower position to an upper position. This g does nothing to the vector index, as we see from (ei)j = g' ik (ek)j. So we don't want to think that somehow ek is a covariant vector and ek a contravariant vector based on the label location. Each of these vectors is in fact both a covariant and a contravariant vector depending on whether we choose to put the index down or up, and we have seen that in fact (ek)j = (ek)j and (ek)j = (ek)j so there is no distinction for these vectors between covariant and contravariant. That is, they are just vectors in Cartesian space. (b) The Reciprocal Base Vectors in 3D. Now, in a slight break from procedure, we want to discuss these reciprocal vectors in 3D only, and then in section (e) below we shall show that the 3D definitions given below for the 3D reciprocal base vectors agree with the general definition given above. In other words, we shall now engage in the traditional presentation of the 3D reciprocal vectors as one finds in various books like M&M. After all, most of our curvilinear coordinate systems of interest will be 3D ones. So here then is a picture (in Cartesian x space) showing our original base vectors ei as well as the new reciprocal ei ones. The 3D reciprocal base vectors are obtained by taking cross products of the tangent base vectors, then dividing by v : e1 ≡ e2 x e3/v e2 ≡ e3 x e1/v e3 ≡ e1 x e2/v or ei = ½ εijk ej x ek / v // remember this when reading section (e) later Now, first recalling this vector identity, (A x B) x (A x D) = [A,B,D] A where [a,b,c] ≡ a (b x c) = etc (see earlier) we can consider the cross product of two of the reciprocal lattice vectors e1 x e2 = (e2 x e3) x (e3 x e1)/v2 = – (e3 x e2) x (e3 x e1)/v2 = - [e3, e2, e1] e3/ v2 = e3/v The results for all three cases are then shown on the first line below, while we copy down the earlier results to the second line for comparison: e1 ≡ e2 x e3 v e2 ≡ e3 x e1 v e3 ≡ e1 x e2 v or ei = ½ εijk ej x ek v e1 ≡ e2 x e3/v e2 ≡ e3 x e1/v e3 ≡ e1 x e2/v or ei = ½ εijk ej x ek / v This similar nature of these two lines is the one justification of the word "reciprocal". In general, neither the ei nor the ei form an orthogonal set, but they are "cross orthonormal" : em en = em (½ εnjk ej x ek / v) = ½ v-1 εnjk [em, ej,ek ] = ½ v-1 εnjkεmjk v = ½ (εnjkεmjk) = 1/2 ( 2δmn) = δmn  This result is a little slippery. One must remember that none of these base vectors is a unit vector. So when we find that e3 e3 = 1, that does not mean (1) e3 and e3 are the same, (2) e3 and e3 are even in the same direction. In fact, e3 and e3 are generally in different directions. It is true that e3 is perpendicular to both e1 and e2 , but that is obvious from the cross product definition of e3. By scaling with 1/v for the reciprocal vectors and v for the tangent vectors, we have contrived to make e3 e3 = 1. Notice that [e1, e2,e3 ] = (e1) (e2x e3) = (e2 x e3/v) (e1/v) = [e1, e2, e3]/ v2 = v/v2 = 1/v So the volume of the reciprocal base vector parallelepiped is 1/v, compared to [e1, e2, e3] = v . (c) A confusing notational issue concerning vectors. This issue is always confusing to the author, so perhaps it is as well to some readers. We should have it out here and be done with it. Let V be a vector in Cartesian space (where up or down indices do not matter) V = (Vx, Vy, Vz) Vx = V Vi = V // so = Our (covariant) transformation is V'a = Tab Vb . Suppose T is a simple rotation. In the active view, V' is some new rotated vector in x-space that differs from V. In the passive view, V' is the same vector as V, and the prime just means it is being viewed in a rotated frame of reference which is x' space. The active view is always easier to understand (for me) since we have V'i = RijVj and V' ≠ V In the passive view, we say this vector vee = V'1 ' + V'2 ' + V'3 ' = V1 + V2 + V3 The notational issue is this: what notation do we use for "vector vee" ? We really have to use the notation vector vee = V. Suppose our x-space system were called the x"-space system. Then the above would say vector vee = V'1 ' + V'2 ' + V'3 ' = V"1 " + V"2 "+ V"3 " and then we would be happy to say V = V'1 ' + V'2 ' + V'3 ' = V"1 " + V"2 "+ V"3 " and we might then add this notation, where we mark the parentheses to show which coordinate system we are talking about. V = (V'1, V'2, V'3)' = (V"1, V"2, V"3)" It just so happens that our Cartesian starting system has no marking like ", it has a blank marking. In this case we have V = (V'1, V'2, V'3)' = (V1, V2, V3) In general, staying with our notation used all along, we could say V = (V'1, V'2, V'3)x'-space = (V1, V2, V3)x-space or V = (V'1, V'2, V'3)q-space = (V1, V2, V3)Cartesian We might go so far as to say V = (V1, V2, V3)x = V1 + V2 + V3 V' = (V'1, V'2, V'3)x' = V'1 ' + V'2 ' + V'3 ' V = V' and this is what we mean by saying that V' is the same vector as V in the passive view. (d) Vectors expanded on the base vectors in various ways. In x space we have V = V*1 + V*2 + V*3 where we add extra * labels to indicate that V*i are the usual Cartesian components of V. Using our passive viewpoint, in q-space we can expand this same vector as We can alternatively write our same vector V as a sum of its projections on either the base vectors or the reciprocal vectors (which in general are not unit vectors) V = V1 e1 + V2 e2 + V3 e2 = Vi ei (*) V = V1 e1 + V2 e2 + V3 e2 = Vi ei (**) where our choice of up and down positions for the indices on the V's will be justified in a moment. So we have three different sets of coordinates for the same vector V. Using our earlier notation, V = (V*1, V*2, V*3)cart V = (V1, V2, V3)base V = (V1, V2, V3)recip Question: Do the three components V1, V2, V3 comprise a contravariant vector? Now go back to our rules for inverse transformations as shown at the end of Section 4 (c), V*a = V 'b Tba V*a = V 'b Tba On the RHS we can replace the T's with (ei)j = Tij and (ei)j = Tij to get V*a = V 'b (eb)a V*a = V 'b (eb)a Everywhere it appears, the index a is a Cartesian space component label. On the LHS, the up and down components are the same since these are Cartesian components. So if we remove the a labels, we get V = V 'b eb V = V 'b eb If we compare these equations with (*) and (**) above, we can identify V 'b = Vb V 'b = Vb Therefore, in our two expansions above, V = V1 e1 + V2 e2 + V3 e2 = Vi ei (*) V = V1 e1 + V2 e2 + V3 e2 = Vi ei (**) we are fully justified in referring to the Vi as contravariant components of V and Vi as covariant components of V, both in q-space. So that explains why we put the up/down sense as we did on these components. And of course we then know that Vi = g'ijVj and Vi = g'ijVj where g' is the metric tensor in q space (= x' space). In Cartesian space we used unit vectors like , but in q-space we have to always keep in mind that none of the 6 base vectors are unit vectors. We could normalize them if we wanted i ≡ ei /|ei| and define vi = |ei| Vi so that Vi ei = vii i ≡ ei /|ei| and define vi = |ei| Vi so that Vi ei = vii Then we can form yet two more expansions of V V = v1 1 + v2 2 + v3 2 = vi i (*) V = v1 1 + v2 2 + v3 2 = vi i (**) So just for fun, we gather up our five different expansions of the same vector V V = V1 e1 + V2 e2 + V3 e2 = Vi ei Vi = g'ijVj V = V1 e1 + V2 e2 + V3 e2 = Vi ei Vi = g'ijVj V = v1 1 + v2 2 + v3 2 = vi i vi = |ei| Vi V = v1 1 + v2 2 + v3 2 = vi i vi = |ei| Vi V = V*1 + V*2 + V*3 = V*i V*i = V j Tji = Vj Tji There is an important warning here. When one talks about the components of vector V, one must be very clear as to the meaning of these components. What basis vectors are we expanding onto? This issue will come up soon in considering the transformation of differential operators like the divergence operator. (e) Verify that both definitions of the ei agree; the generalized cross product We shall first prove a certain relationship between T matrices [the multi-T relation (6) below] , then we shall use that to show that, for N=3 , our general definition of the ei matches the geometric cross product definition used above. To get this "multi T" relation, we have to thread a few needles, so here we go. Start with this known fact, v = det(Tdn) = (εjst.... Tkj Tas Tbt.... ) if k,a,b.... are all different (1) Rewrite (1) in the following way, which is valid for any values of k,a,b...., εkab.... det(Tdn) = (εjst.... Tkj Tas Tbt..... ) (2) Notice that both sides are 0 if two indices from the kab... set are the same, since then the determinant will have two identical rows. We will next make use of this fact, εiab... εkab... = (N-1)! δik, (3) which perhaps requires a little explanation. Pick a value of i and consider one term in the sum on ab... Then for the first ε to be non-zero, abc... must be a permutation of all the other N-1 integers between 1 and N. This same index set appears on the second ε. If k ≠ i, the integer k will appear twice so second ε = 0 and then LHS = 0, matching the RHS. On the other hand, if k = i, there are (N-1)! ways to select the integer set abc... and each way produces either (+)(+) or (-1)(-1), giving a sum of (N-1)!. The familiar version of this sum rule in 3D is εiab εkab = 2δik. Now multiply both sides of (1) by δik using the fact (3) on the RHS to get v δik = [1/(N-1)!] εiab... { εkab... det(Tdn)} (4) Insert (2) into (4) to replace {..} , v δik = [1/(N-1)!] εiab... (εjst... Tkj Tas Tbt... ) Apply Tkc to both sides and sum on k v Tkc δik = [1/(N-1)!] εiab... (εjst... Tkc Tkj Tas Tbt... ) Use orthogonality on the RHS so that Tkc Tkj = δcj , vTkc δik = [1/(N-1)!] εiab... (εjst... δjc Tas Tbt... ) or vTic = [1/(N-1)!] εiab... εcst...Tas Tbt... or Tij = [1/(N-1)!] εiab... εjst...Tas Tbt.... / v // multi-T formula (5) which is our "multi T relation" which says that a down-tilt T matrix element is a linear combination of products of N-1 up-tilt matrix elements, divided by v = det(T). Now we replace the T on the left with (ex)y ≡ Txy and the T's on the right with (ex)y ≡ Txy, (ei)j = [1/(N-1)!] εiab... εjst...(ea)s (eb)t... /v (6) So we can express any reciprocal base vector in terms of the tangent base vectors. Equation (6) is a generalized cross product of the tangent base vectors divided by (N-1)! v . For N=3 it becomes (ei)j = (1/2) εiab εjst(ea)s (eb)t/v = (1/2) εiab [ea x eb]j /v or ei = (1/2) εiab ea x eb / v which matches the 3D definition of the reciprocal vectors we used above. Of course this can also be written e1 = e2 x e3 / v e2 = e3 x e1 / v e3 = e1 x e2 / v We can write (6) in the following more suggestive notation ei = [1/(N-1)!] εiab... ea x eb x ec ... /v (6') an example of which would be e1 = e2 x e3 x e4 ... x eN / v [ I would call (6) a "super generalized cross product" since it is summed on 2(N-1) indices! If we take (6) and just sum on the abc... indices, we generate (N-1)! identical terms when the vectors are put into the same order as the first term times (-1)i-1 (hint of this below). Then we get a simpler result: (ei)j = (-1)i-1εjst...(e1)s (e2)t... /v (6A) and this is what I would call the "generalized cross product", and it agrees with (A.2) from Appendix A of my new tensor doc. Here I have not clarified that there are only N-1 en factors and that ei is missing from the right!! We do have v = detS which works right. ] To justify this notation, we consider a more generic case of a generalized cross product of N-1 contravariant vectors each of dimension N to form a covariant vector of dimension N, Q = B x C x .... x X // cross product of N-1 vectors Qa = εabc..x BbCc...Xx // this line defines the notation on the previous line The notation with the crosses suggests that Q is orthogonal to the vectors of which it is composed, and this is in fact true. For example Q.C = QaCa = εabc..x BbCc...Xx Ca = Σcd..x [ Σab εabc..x CcCa] BbDd... Xx = 0 because the [..] object vanishes since we have a symmetric tensor contracted with two indices of a totally antisymmetric tensor. Looking at (6'), we see then that ei is orthogonal to all the ej with j ≠i, something we already knew, since ei. ej = ei ej = δij. The cross product also suggests, based on the N=3 idea, that if we swap two vectors, the sign of the cross product will change. This is in fact also true, and here is an example: Q = B x C x D x E x F Qa = εabcdefBbCcDdEeFf We now swap indices c and e on ε, creating a minus sign, and then swap positions of Cc and Ee = – εabedcfBbCcDdEeFf = - εabedcfBbEeDdCcFf and finally we swap the names of dummy summation variables e↔c = - εabcdefBbEcDdCdFf = – [ B x E x D x C x F ]a And therefore, B x C x D x E x F = – B x E x D x C x F and it should be clear that this idea works no matter which pair of vectors we swap. 12. The transformation under T of differential distance, area and volume elements. (a) Distance. (general N) From our definition of ei , which is that (ei)j = [ ∂x/∂qi]j , we have dx = ei dqi . We can then reverse the tilt based on our tilt rule discussed at the end of section (a), so dx = ei dqi = ei dqi dxj = (ei)j dqi and dxj = (ei)j dqi (ei)j = Tij = [ ∂x/∂qi]j (ei)j = Tij Similarly, from our definition of e'i , which is that (e'i)j = [ ∂q/∂xi]j , we have dq = e'i dxi so dq = e'i dxi = e'i dxi dqj = (e'i)j dxi and dqj = (e'i)j dxi (e'i)j = Tji = [ ∂q/∂xi]j = (T-1)ij (e'i)j = Tji = (T-1)ij So given any linear change vector dx in x space (either contravariant of covariant), we know the dq it causes in q-space. So this is how differential linear things transform. In terms of the scalar distance we have ds2 = gijdxidxj = dxidxi = (dx1)2 + (dx2)2 + (dx3)2 = dx dx x-space ds2 = g'ijdqidqj = dqidqi = dq.dq q-space We know ds2 is the same in both lines since it is a tensorial scalar under T. (b) Area. (N=3 only) We make these three definitions, each corresponding by an individual variation dqi, ds1 = e1 dq1 ds2 = e2 dq2 ds3 = e3 dq3 These dsi are distances in the Cartesian x space. We can write dsi = ei dqi where we have a rare instance of a tilted index NOT being summed. Now define these differential vector 2D areas: dS1 = ds2 x ds3 dS2 = ds3 x ds1 dS3 = ds1 x ds2 or dSi = ½ εijk dsj x dsk = ½ εijk [ej x ek ] dqj dqk = ½ εijk [ej ek ] | ej x ek | dqj dqk where [ej ek ] is meant to be a unit vector in the direction ej x ek. What we care about is the magnitude of this vector area, so we have dSi = ½ εijk | ej x ek | dqj dqk We go off then and grab our vector identity, (A x B) (A x B) = A2B2 – (AB)2 so that | ej x ek | 2 = (ej x ek) (ej x ek) = (ejej) (ekek) - (ejek)2 = g'jjg'kk - g'jk2 giving the result dSi = ½ εijk dqj dqk for example dS1 = dq2 dq3 and cyclic This rule can be generalized to N dimensions using the generalized cross product formula derived in Section 11 (e), but we shall not attempt that here. [ This generalization was what got me writing the new doc. I did not realize in this old doc that the above dS1 was a cofactor of g'11. ] [ Below is my first attempt to generalize the above to N, added well after the original doc was written (added perhaps Oct 1 2011).] (b1) Area. (general N and N=3) In N dimensions for N≠3, the notation q = a x b has no meaning so we must replace it with something that does have meaning. Consider the N=3 case again for two vectors called a(1) and a(2). The cross product is written this way (repeated indices are summed unless stated otherwise) q = a(1) x a(2) q = A x B qk = εkii a(1)i a(2)i qk = εkab Aa Bb The generalization to N dimensions is this (there must be N-1 vectors crossed together, so there will then be N-2 cross's ) q = a(1) x a(2) x a(3) x .... x a(N-1) q = A x B x C x ... x X qk = εkiii....i a(1)i a(2)i .... a(N-1)i qk = εkabc..x Aa BbCc ... Xx where this ε is the totally antisymmetric tensor in N dimensions (it has N indices), meaning that if any two indices are swapped, it changes sign. It is easy to show that this N-vector q is perpendicular to all the a(i) , just as we are used to having when N = 3: a(2) q = a(2)k qk = a(2)k εkiii....i a(1)i a(2)i .... a(N-1)i = εkiii....i a(1)i [ a(2)k a(2)i ] .... a(N-1)i In terms just of indices k and i2 and their sums, we can think of ε as being Aki , an Antisymmetric tensor in these indices. Then we have Aki[ a(2)k a(2)i ] = – Aik[ a(2)k a(2)i ] = – Aki[ a(2)k a(2)i ] and something that is minus itself must be 0. In the last step we just swapped the names of the dummy dummation indices. In general, AijSij = 0 if A is antisymmetric and S is symmetric. Therefore, we have shown that a(2) q = 0, and in general q is perpendicular to all the a(m) which appear in the cross product. We are then led to this obvious generalization of the formula for differential area in N dimensions dS1 = ds2 x ds3 x ds4 .... x dsN where as desired this area element is perpendicular to all the vector distances dsk from which it is constructed. The general case can be written dSk = Πx;i≠k(x) dsi where the x indicates the mulitplicative operation indended by Π. As for N=3, the vectors dsi skaffold an N dimensional parallelogram in N-space, and each of these vectors can be written in terms of its tangent base vector ei dsi = ei dqi i = 1...N so we can write the area element this way dSk = Πx;i≠k(x) dsi = (Πx;i≠k ei) (Πi≠k dqi) For example, in N=5 dimensions we would have dS3 = e1 x e2 x e4 x e5 dq1dq2dq4dq5 We are interested in the magnitude dSk so we have to compute the magnitude2 of (Πx;i≠k ei ): (Πx;i≠k ei ) (Πx;i≠k ei) = (Πx;i≠k ei )j (Πx;i≠k ei )j Now using our cross product ε formula from above we can write (Πx;i≠k ei )j = εjiii....i (e1)i (e2)i .... (eN)i where ε is missing index ik and the product is missing (ek)i ; the reader will just have to keep this in mind so we don't have to show it in the notation. Then we have (Πx;i≠k ei )j (Πx;i≠k ei )j = (εjiii....i (e1)i (e2)i .... (eN)i )(εjmmm....m (e1)m (e2)m .... (eN)m) = εjiii....i εjmmm....m [(e1)i(e1)m][(e2)i(e2)m] ..... [(eN)i(eN)m] Now we have to deal with this εε product. So we pause to consider: Each of the "other terms" has the same form as the first term, but the set of mi labels has been permuted (and the ii labels stay fixed). Such a term has a sign ± depending on whether this permutation is obtained by an even or odd number of index swaps from the original order. Including the first term, the total number of permutations is (N-1)!, so the portion "other terms" contains (N-1)! - 1 terms. For N=3, there is only one "other term". Notice that no mi in any term (including the first term) ever takes the value k needs repair Therefore, we have shown that (Πx;i≠k ei ) (Πx;i≠k ei) = [e1 e1] [e2 e2]..... [eN eN] + other terms where again [ek ek] is missing. In those other terms, the second set of e's is permuted. We recognize that dot products as metric tensor elements, so we have (Πx;i≠k ei ) (Πx;i≠k ei) = g'11 g'22..... g'NN + other terms Now that we have a relatively simple product, we can account for the signs due to the swaps mentioned above with and ε symbol having N-1 indices. Thus (Πx;i≠k ei ) (Πx;i≠k ei) = g'1m g'2m..... g'Nm εmmm....m // mk missing For example, if N=4 and k=2 then each index mi in the implied sums only take the values mi = 1,3,4 because we can never have In the implied sums here over the mi, we are listing off the permutations of {mi...} = { 1,2....N} where k is missing from the list. Therefore the sums are all of the form Σm≠m . We can add this reminder to the notation by writing (Πx;i≠k ei ) (Πx;i≠k ei) = Σm≠m g'1m g'2m..... g'Nm εmmm....m // mk missing So there are two issues here. (1) ε is missing index mk and g'km is missing from the product of g' objects. (2) each sum excludes mk We have therefore shown that dSk = Πx;i≠k(x) dsi = (Πx;i≠k ei) dSk = [Σm≠m g'1m g'2m..... g'Nm εjmmm....m]1/2(Πi≠k dqi) // mk missing As a check on this result, for N=2 it says dS1 = [Σm≠1g'2m g'3m εmm]1/2 dq2 dq3 = [g'2m g'3m εmm]1/2 dq2 dq3 g'2m g'3m εmm [ I did not realize that I had a cofactor thing here and just gave up. This is done correctly now in the new doc in Section 8 (d). ] (c) Volume. (N=3 and general N) We can write the 3D volume element as follows dV = ds1 dS1 = ds1 ds2 x ds3 = e1 e2 x e3 dq1 dq2 dq3 = v dq1 dq2 dq3 so v is the differential volume transformation factor, dV = v dq1 dq2 dq3 = dx1 dx2 dx3 where the RHS is the Cartesian volume. But we have already proven the general result for N dimensions in Section 10 (c) above, which was this, [ qi = x'i everywhere ] [dx1dx2dx3...... ] = v [ dq1dq2dq3..... ] v = det(Tab) = det () = det J(x,x') 1/v = det(Tab) = det() = det J(x',x) v = det(Tij) = det(Tup) = det [ (e1) , (e2), (e3), ..... ] = εijk.... (e1)i(e2)j (e3)k ..... = εijk.... T1i T2j T3k ..... where J is called the Jacobian Matrix. We have already shown in Section XXX that v2 = det(g'), so we now have the following entertaining relationships (all valid for N dimensions) v = factor by which the differential volume transforms under T v = volume of the parallelepiped ( e1 e2 x e3 in 3D ) v = where g' is the metric tensor (in q-space) v = det(Tab) = det( ) = J(x,x') = the Jacobian (d) Summary of the results. dx = ei dqi = ei dqi dxj = (ei)j dqi and dxj = (ei)j dqi (ei)j = Tij = [ ∂x/∂qi]j (ei)j = Tij dq = e'i dxi = e'i dxi dqj = (e'i)j dxi and dqj = (e'i)j dxi (e'i)j = Tji = [ ∂q/∂xi]j = (T-1)ij (e'i)j = Tji = (T-1)ij ds2 = g'ijdqidqj = dxidxi dsi ≡ eidqi // no sum dS1 = ds2 x ds3 and cyclic // 3D dS1 = dq2 dq3 = dx2dx3 and cyclic // 3D dV = dq1 dq2 dq3... = dx1 dx2 dx3... ±v = det(g') = det(Tab) = det( ) = J(x,x') // = ( e1 e2 x e3 ) in 3D ±1/v = det(Tab) = det() = det J(x',x) // = ( e1 e2 x e3 ) in 3D 13. The Divergence in curvilinear coordinates. (a) Definition of div A. Let's go back to our parallelepiped drawings, which recall are drawn in Cartesian x-space, We actually want do deal with a differential scaled-down version of the above parallelepiped which has these edges ds1 = e1 dq1 ds2 = e2 dq2 ds3 = e3 dq3 We consider a vector field A(x) where x = x(q) under some curvilinear transformation. The divergence of A is then defined with respect to this differential parallelepiped as ( we dispense with limit notation here, it is implied) div A ≡ ∫AdS /dV (*) where dV is volume of the parallelepiped, which we know to be dV = |ds1 (ds2 x ds3)| = |e1 (e2 x e3)| dq1 dq2 dq3 = v dq1 dq2 dq3 (b) Calculation of the surface integral. The entire calculation will be done in Cartesian space. The face nearest the viewer we shall call the "front" face. The outward pointing vector area of this front face is dSfront = ds1 x ds3 = (e1 dq1) x ( e3 dq3) = [ e1 x e3 ] dq1 dq3 = - e2 v dq1 dq3 where e2 is a reciprocal base vector, and v is the parallelepiped volume, both evaluated at the center of the front face. Remember that both these objects change as we move away from x = f(q) as per our earlier discussion. We are then interested in A dSfront = - (A e2) v dq1 dq3 But (A e2) = A2, the contravariant component of vector A ( since (Aiei) e2 = A2), so we have A dSfront = - A2(front) v (front) dq1 dq3 Similarly we find that A dSback = +A2(back) v(back) dq1 dq3 where the sign change arises since the outgoing vector area on the back points oppositely to the outgoing vector area on the front. Also, we have to evaluate A2 and v at the center of the back face. Adding these two, we find A dSfront + A dSback = - [ (v A2) |front - (v A2) |back ] dq1 dq3 = - [ (v A2)(q) - (v A2)(q + dq2 ) ] dq1 dq3 = [ (v A2)(q + dq2) - (v A2)(q) ] dq1 dq3 = [ ∂q2 (v A2) dq2] dq1 dq3 = [ ∂q2 (v A2) ] dq1 dq2 dq3 where dq2 means dq2 2 and we are treating (v A2)(q) as a function like h(q). If we treat the other two pairs of faces in the same way, we shall find that ∫AdS = [ ∂q2 (v A2) + ∂q1 (v A1) + ∂q3 (v A3) ] dq1 dq2 dq3 = ∂qi (vAi) dq1 dq2 dq3 (c) Solve for div A Inserting our computed surface integral into our definition of div A shown in (*) above, div A ≡ ∫AdS /dV dV = v dq1 dq2 dq3 we get div A = ∂qi (vAi) dq1 dq2 dq3 / ( v dq1 dq2 dq3 ) = (1/v) ∂qi (vAi) A = A1e2 + A2e2 + A3e3 We now introduce a traditional symbol "g' " (still maintaining our prime since this is in x'-space) g' ≡ det(g') => v = // as shown in Section 10 (e) so we restate our result in the traditional manner div A = (1/) ∂i (Ai) g' ≡ det(g') A = A1e1 + A2e2 + A3e3 where it is understood that all objects are a function of the curvilinear coordinates q, so ∂i = ∂qi . (c) Various forms of div A. We have to be very careful about the meaning of the components of vector A. Recall how earlier at the end of section 11 (d) we showed five different ways to define the components of A. In our work above, we have made use of a particular one of these five ways -- the one shown above, where the ei are not unit vectors. Usually, in practice, one wants to work with a set of components which use unit vectors, namely. A = a1 1 + a2 2 + a3 2 = ai i ai = |ei| Ai = QiAi where the last equality follows from |ei| = = = Qi => ei = i Qi , We can then express our divergence result as div A = (1/) ∂i (ai/Qi) A = a1 1 + a2 2 + a3 2 But of course no one would write it this way, using two different notations for the same vector, so at this point we replace the ai with Ai (not the same as the previous Ai) to get div A = (1/) ∂qi (Ai/Qi) A = A1 1 + A2 2 + A3 2 (c) div A in orthogonal systems. Now, if g' is diagonal so we have an orthogonal system, then making the usual definition Qi2 = g'ii we get = = Q1Q2Q3 and we get this result div A = (1/[ Q1Q2Q3] ) ∂i (Q1Q2Q3 Ai /Qi) = (1/[ Q1Q2Q3] ) [ ∂1 (Q2Q3 A1) + cyclic ] // orthog, A = A1 1 + A2 2 + A3 2 Although we have been working in 3D above, it seems clear that the result in N dimensions is this : div A = (1/) ∂i (Ai) where A = Aiei g = det(g') div A = (1/) ∂i (Ai/Qi) where A = Ai i For example, for N=4, our surface integral would involve 4 pairs of faces instead of 3, and v would be the volume of the 4D parallelepiped. 14. The Gradient in curvilinear coordinates. (a) Definition of grad(f) The gradient is defined as follows: ( I decided not to "bold" the letters grad ) df = f(x + dx) - f(x) = grad(f) dx = F dx (*) where we make up a convenient shorthand symbol F = grad(f) . Our curvilinear transformation gives us some x = x(q) so we may regard f as a function of q. (b) Calculation of grad(f) Recall from earlier that we have ds1 = e1 dq1 ds2 = e2 dq2 ds3 = e3 dq3 for our candidate displacements dx. Since we want to solve for F, we expand it with covariant coefficients, since we know that ei ej = δij, F = Fiei and we pick dx = ds1 . We then get F dx = (Fiei) (e1 dq1) = F1 dq1 On the other hand, we know that in this case df = (∂q1f)dq1 Therefore we find from (*), perhaps to no great surprise, since we know ∂i transforms as a covariant vector, F1 = (∂q1f) and similarly for the F2 and F3 components. Thus F = Fiei = (∂qif) ei // sum on i (c) Various forms of grad(f) We usually want this expanded on the e1 tangent base vectors and not on the reciprocal ones, so F = (∂qif) g'ij ej If we want this expanded on unit tangent base vectors then using |ei| = = = Qi => ei = i Qi we find that F = [(∂qif) Qj g'ij] j and in the case of orthogonal coordinates, g'ij = Qi-2 δij [ remember g'ab is the inverse of g'ab] so this becomes F = [(1/Qi) (∂qif)] i Replace F by grad(f) , and replace ∂qi = ∂i with the understanding of the meaning, to get: grad(f) = (∂if) ei = (∂if) gij ej = [(∂if) Qj g'ij] j // general In an orthogonal system this becomes grad(f) = (∂if) ei = (1/Qi)2(∂if) ei = (1/Qi) (∂if) i // orthog The work above, though stated for 3D, is clearly valid in N dimensions. (d) The combination A grad(f) Going back to our F notation just a bit more, sometimes we are interested in this combination A F = A grad(f) As usual, we must take care with the meaning of the coefficients of A. Here are some cases: A = Akek => A F = (Akek) [(∂if) gij ej ] = Aj gji (∂if) // general = Ai (1/Qi)2(∂if) // orthog A = Akek => A F = (Akek) (∂qif) ei = Ai (∂qif) // general A = Akk => A F = (Akk) (∂qif) ei = Ai (1/Qi) (∂qif) // general where we have not bothered with the intermediary ak notation, so the Ak are different animals on these last two lines. Write these out one more time, A grad(f) = Aj gji (∂if) A = Akek // general = Ai (1/Qi)2(∂if) // orthog A grad(f) = Ai (∂if) A = Akek // general A grad(f) = (Ai/Qi) (∂if) A = Akk // general where the last line is the "practical" form. Sometimes people are not expecting the Qi sitting there. 15. The Laplacian in curvilinear coordinates. The Laplacian is defined by (we continue to use F as a shorthand for grad(f) ) lap(f) = div grad(f) = div F and in the previous two sections we have assembled the required ingredients. First, we know that div F = (1/) ∂i (Fi) where F = F1e1 g' = det(g') Second, we know that Fi = (∂if) so Fi = gij (∂jf) and we end up with lap(f) = div grad(f) = (1/) ∂i [g'ij (∂jf) ] // general where now we have no issues with how components of vectors are selected. In orthogonal coordinates we have = = Q1Q2Q3 g'ij = Qi-2 δij so lap(f) = div grad(f) = (1/ [Q1Q2Q3]) ∂i [Q1Q2Q3 (Qi)-2 (∂if) ] // orthog = (1/ [Q1Q2Q3]) { ∂1 [ (Q2Q3/Q1 ) (∂1f)] + cyclic } = (1/Q1)2 ∂12f + {[∂1 (Q2Q3/Q1 )]/ [Q1Q2Q3]} (∂1f) + cyclic 16. The Curl in curvilinear coordinates. (a) Definition of curl A. As usual, we consider the scaled down differential version of our parallelepiped which has edges ds1 = e1 dq1 ds2 = e2 dq2 ds3 = e3 dq3 and we consider a vector field A(x) where x = x(q) under some curvilinear transformation. The quantity curl A is a vector whose ith component is the "circulation integral of A" around "positive" face i of the parallelepiped, divided by the area of that face. A face is "positive" if its area vector points "out". For example, the top face in our picture has area dS3 = ds1 x ds2 = dq1dq2 e1 x e2 = dq1dq2 v e3 which points up and therefore "out". When we say " ith component" here, it is implied that this is a component of the curl when the curl is expanded on unit vectors. So, according to this definition, we have (curl A) i = ( Ads)i / dSi (curl A) dSi i = ( Ads)i (curl A) dSi = ( Ads)i As we did with the gradient, we invent a temporary name for the curl, F = curl A, and then replace F at the end with curl A. So we want to compute F in curvilinear coordinates, where F satisfies this equation F dSi = ( Ads)i (*) (b) Computation of the circulation. We shall do this for the top face of our "scaled down" parallelepiped The circulation is a counterclockwise (CCW) line integral around the boundary of the top face. Thus, ( Ads)3 = [A(front) - A(back)] ds1 + [A(right) - A(left)] ds2 = [A(front) - A(back)] (e1 dq1) + [A(right) - A(left)] (e2 dq2) where A(front) refers to the value of A at the center of the "front" edge of the parallelogram which is the top face, and similarly for the other three edge centers. Motivated by ei ej = δij, we expand A as follows A = Ajej // sum on j which then gives us ( Ads)3 = [A1(front) - A1(back)] dq1 + [A2(right) - A2(left)] dq2 But we know that A1(back) = A1(front) + (∂A1/∂q2) dq2 A2(right) = A2(left) + (∂A2/∂q1) dq1 so our circulation integral is then ( Ads)3 = [ – (∂A1/∂q2) dq2] dq1 + [(∂A2/∂q1) dq1] dq2 = { (∂A2/∂q1) – (∂A1/∂q2) } dq1 dq2 = { ∂1A2 – ∂2A1} dq1 dq2 (c) Solving for curl A. Installing the above in equation (*) gives F dSi = ( Ads)i = { ∂1A2 – ∂2A1} dq1 dq2 Recalling that dS3 = ds1 x ds2 = dq1dq2 [ e1 x e2] = dq1dq2 v e3 we get that v F e3 = { ∂1 A2 – ∂2A1} Motivated once again by ei ej = δij, we expand F as follows F = Fiei so our result is then F3 = (1/v) { ∂1 A2 – ∂2A1} = (1/v)ε3jk∂j Ak If we analyze the other two positive surfaces in the same manner, we get similar results, which can then be gathered into this single compact statement (where of course ∂j means ∂qj ) Fi = (1/v)εijk∂j Ak A = Aiei F = Fiei (d) various forms of curl A . Usually, we want the contravariant components of A to appear, so write this as [ and recall v = ] Fi = (1/)εijk∂j (g 'knAn) A = Aiei F = Fiei Now if one wants both vectors expanded on unit vectors, recalling that |ei| = = = Qi => ei = i Qi , we rescale the components so that A = Aiei = ai i = aiei/Qi so Ai = ai/Qi, and similarly for F, to get fi/Qi = (1/)εijk∂j (g 'knan/Qn) A = aii F = fii but then we redefine the names of the components to be capital letters to get Fi/Qi = (1/)εijk∂j (g 'knAn/Qn) A = Aii F = Fii or Fi = (Qi/)εijk∂j (g 'knAn/Qn) A = Aii F = Fii So now let's replace F = curl A and write out our results both ways: [curl A] i = (1/)εijk∂j (g 'knAn) A = Aiei curl A = [curl A]i ei => curl A = (1/) ei εijk∂j (g 'knAn) g' ≡ det(g') [curl A] i = (Qi/)εijk∂j (g 'knAn/Qn) A = Aii curl A = [curl A]i i => curl A = (1/) Qi i εijk∂j (g 'knAn/Qn) g' ≡ det(g') = (1/) εijk [Qi i] [ ∂j] [(g 'knAn/Qn) ] = (1/) εijk [M1i ][ M2j] [M3k] = (1/)det(M) where in the last two lines we have defined a matrix whose determinant gives us what we want M = This form, perhaps interesting looking, is just a way to display the result above. (e) curl A in orthogonal systems. For an orthogonal system, we make the usual simplifications g 'knAn = g 'kkAk = Qk2Ak // no k sum g 'knAn/Qn = g 'kkAk/Qk = (Qk)2Ak/Qk = QkAk // no k sum g' = det(g') = Q12 Q22 Q32 so our two results become ORTHOGONAL: [curl A] i = (1/[Q1Q2Q3]) εijk∂j (Qk2Ak) A = Aiei curl A = [curl A]i ei => curl A = (1/[Q1Q2Q3]) ei εijk∂j (Qk2Ak) g' ≡ det(g') ____________________________________________________________________________________ [curl A] i = (Qi /) εijk∂j (QkAk) A = Aii curl A = [curl A]i i => curl A = ( Qi/ [Q1Q2Q3] )i εijk∂j (QkAk) g' ≡ det(g') or [curl A] 1 = (1 /[Q2Q3]) [ ∂2 (Q3A3) – ∂3 (Q2A2) ] and cyclic We shall not attempt to generalize the curl to a number of dimensions different from N=3. 17. The Vector Laplacian in curvilinear coordinates. The definition of this vector object is this: 2A ≡ grad(div A) – curl (curl A) In Cartesian coordinates, one finds that [2A]i = 2(Ai) but in general curvilinear coordinates it's a nightmare. We do have all the necessary pieces from our previous sections, so all we have to do is assemble them. We shall always assume vectors are expanded on the unit vectors i . Let's define f = div A F = curl A so 2A = grad(f) - curl F We start with our earlier expression for F = curl A, F = curl A = (1/) Qi i εijk∂j (g 'knAn/Qn) Fi = [curl A] i = (Qi/)εijk∂j (g 'knAn/Qn) Fn = [curl A] n = (Qn/)εnrs∂r (g 'stAt/Qt) Then insert Fn into the expression for curl F, curl F = (1/) Qi i εijk∂j (g 'kn [Fn]/Qn) = (1/) Qi i εijk∂j (g 'kn(Qn/)εnrs∂r (g 'stAt/Qt) /Qn) = (1/) Qi εijk∂j [g 'kn(1/)εnrs∂r (g 'stAt/Qt) ] i Then for the grad(f) term we have f = div A = (1/) ∂k (Ak/Qk) grad(f) = (∂if) Qi g'ij j = (∂jf) Qj g'ji i // swap i↔j = (∂j[ (1/) ∂k (Ak/Qk) ] ) Qj g'ji i So our complete result is then 2A = (∂j[ (1/) ∂k (Ak/Qk) ] ) Qj g'ji i // general result - (1/) Qi εijk∂j [g 'kn(1/)εnrs∂r (g 'stAt/Qt) ] i where the first term has 3 and the second term 7 summation indices! For orthogonal systems, it is still quite a mess, We insert g'ji = δij Qi-2 etc, to get 2A = (∂i[ (1/) ∂k (Ak/Qk) ] ) (1/Qi) i - (1/) Qi εijk∂j [Qk2(1/)εkrs∂r (Qs2As/Qs) ] i or 2A = (1/Qi) ∂i[ (1/) ∂k (Ak/Qk) ] i - (1/) Qi εijk ∂j [Qk2 (1/)εkrs∂r (QsAs) ] i or 2A = (1/Qi) ∂i[ (1/) ∂k (Ak/Qk) ] i - (1/) Qi εkij εkrs ∂j [Qk2 (1/) ∂r (QsAs) ] i At this point we can use εkij εkrs = δirδjs - δisδjr [ wrong! ] so the second term breaks into two terms 2nd term = - (1/) Qi (δirδjs - δisδjr) ∂j [Qk2 (1/) ∂r (QsAs) ] i = - (1/) Qi∂j [Qk2 (1/) ∂i (QjAj) ] i + (1/) Qi ∂j [Qk2 (1/) ∂j (QiAi) ] i so we then have 3 terms for 2A : 2A = (1/Qi) ∂i[ (1/) ∂k (Ak/Qk) ] i // orthogonal - (1/) Qi∂j [Qk2 (1/) ∂i (QjAj) ] i + (1/) Qi ∂j [Qk2 (1/) ∂j (QiAi) ] i = Q1 Q2 Q3 The first term has 2 summation indices and each of the other two terms has 3, so the number of terms present is 32 + 2*33 = 63, though we might expect some terms to combine. This is a job for Maple. Results for specific orthogonal systems are quoted in M&F and elsewhere. Here is the M&F result for spherical coordinates: ( 1953 vol I p 116 ) where ar = = 3 and so on. Since each scalar Laplacian represents 3 terms, this result contains 18 terms. 18. Summary of the differential operators in curvilinear coordinates. In all these results we have unit vector components A = Aii and g' ≡ det(g') Qi2 = g'ii. (a) for general systems All results are valid for N dimensions except for curl A and 2A A ≡ div A = (1/) ∂i (Ai/Qi) f ≡ grad(f) = (∂if) ei = (∂if) gij ej = [(∂if) Qj g'ij] j A f = A grad(f) = (Ai/Qi) (∂if) 2f ≡ lap(f) = div grad(f) = (1/) ∂i [g'ij (∂jf) ] x A ≡ curl A = (1/) εijk Qii ∂j (g 'knAn/Qn) N = 3 2A = (∂j[ (1/) ∂k (Ak/Qk) ] ) Qj g'ji i - (1/) Qi εijk∂j [g 'kn(1/)εnrs∂r (g 'stAt/Qt) ] i N = 3 (b) for 3D orthogonal systems A ≡ div A = (1/[ Q1Q2Q3] ) [ ∂1 (Q2Q3 A1) + cyclic ] f ≡ grad(f) = (1/Qi) (∂if) i A f = A grad(f) = (Ai/Qi) (∂if) 2f ≡ lap(f) = (1/ [Q1Q2Q3]) { ∂1 [ (Q2Q3/Q1 ) (∂1f)] + cyclic } x A ≡ curl A = (1 / [Q1Q2Q3] ) εijk Qii ∂j (QkAk) or [curl A]1 = (1 /[Q2Q3]) [ ∂2 (Q3A3) – ∂3 (Q2A2) ] and cyclic 2A = (1/Qi) ∂i[ (1/) ∂k (Ak/Qk) ] i - (1/) Qi∂j [Qk2 (1/) ∂i (QjAj) ] i + (1/) Qi ∂j [Qk2 (1/) ∂j (QiAi) ] i = Q1 Q2 Q3 (c) Continued Example: Spherical Coordinates. As we found in Section 10 (f), 1 = 2 = 3 = Q1 = r Q2 = rsinθ Q3 = 1 // often called hθ,hφ, hr so we turn the crank, giving a summary at the end: A ≡ div A = (1/[ Q1Q2Q3] ) { ∂1 (Q2Q3 A1) + ∂2 (Q3Q1 A2) + ∂3 (Q1Q2 A3) } = (1/[r2sinθ] ) { ∂θ (rsinθ Aθ) + ∂φ(r Aφ) + ∂r (r2sinθ Ar) } = (1/[r2sinθ] ) { r ∂θ (sinθ Aθ) + r ∂φ(Aφ) + sinθ ∂r (r2 Ar) } = { (1/rsinθ) ∂θ (sinθ Aθ) + (1/rsinθ) ∂φ(Aφ) + (1/r2)∂r(r2Ar) } f ≡ grad(f) = (1/Qi) (∂if) i = ( 1/r) (∂θf) + (1/rsinθ) (∂φf) + (∂rf) A f = Aθ(1/r) (∂θf) + Aφ(1/rsinθ) (∂φf) + Ar(∂rf) A = Akk 2f = (1/ [Q1Q2Q3]) { ∂1 [ (Q2Q3/Q1 ) (∂1f)] + ∂2 [Q3Q1/Q2 ) (∂2f)] + ∂3 [ (Q1Q2/Q3 ) (∂3f)] } = (1/[r2sinθ] ){ ∂θ [ sinθ (∂θf) ] + ∂φ [ (1/sinθ) (∂φf)] + ∂r [ (r2sinθ) (∂rf)] = (1/[r2sinθ] ){ ∂θ [ sinθ (∂θf) ] + (1/sinθ) ∂φ [(∂φf)] + sinθ ∂r [r2(∂rf)] = (1/r2sinθ))∂θ [ sinθ (∂θf) ] + (1/r2sin2θ) ∂φ [(∂φf)] + (1/r2) ∂r [r2(∂rf)] [curl A]θ = (1 /[Q2Q3]) [ ∂2 (Q3A3) – ∂3 (Q2A2) ] = (1/rsinθ) [ ∂φ (Ar) – ∂r (rsinθAφ) ] = (1/rsinθ) [ ∂φ (Ar) – sinθ ∂r (rAφ) ] = (1/r) [(1/sinθ)∂φ (Ar) – ∂r (rAφ) ] [curl A]φ = (1 /[Q3Q1]) [ ∂3 (Q1A1) – ∂1 (Q3A3) ] = (1/r) [ ∂r (rAθ) – ∂θ (Ar) ] [curl A]r = (1 /[Q1Q2]) [ ∂1 (Q2A2) – ∂2 (Q1A1) ] = (1 /r2sinθ) [ ∂θ (rsinθAφ) – ∂φ (rAθ) ] = (1/rsinθ) [∂θ (sinθAφ) – ∂φ (Aθ) ] Assembling the three pieces: curl A = (1/r) [(1/sinθ)∂φ (Ar) – ∂r (rAφ) ] + (1/r) [ ∂r (rAθ) – ∂θ (Ar) ] + (1/rsinθ) [∂θ (sinθAφ) – ∂φ (Aθ) ] So now we collect the above items and put then in traditional r,θ,φ order: A ≡ div A = (1/r2)∂r(r2Ar) + (1/rsinθ) ∂θ (sinθ Aθ) + (1/rsinθ) (∂φAφ) f ≡ grad(f) = (∂rf) + (1/r) (∂θf) + (1/rsinθ) (∂φf) A f = Ar(∂rf) + Aθ(1/r) (∂θf) + Aφ(1/rsinθ) (∂φf) A = Akk 2f = (1/r2) ∂r [r2(∂rf)] + (1/r2sinθ) ∂θ [ sinθ (∂θf) ] + (1/r2sin2θ) ∂φ2f curl A = (1/rsinθ) [∂θ (sinθAφ) – (∂φAθ) ] + (1/r) [(1/sinθ)(∂φAr) – ∂r(rAφ) ] + (1/r) [ ∂r (rAθ) – (∂θAr) ] The result for the vector Laplacian is quoted at the end of Section 17. ******************************************************************************** Appendix A. Matrix notation. Here we look in more detail at the generic "matrix notation" used in Section 4 (b) and clarify items that might seem hazy such as the meaning of (T-1)up . If we just refer to "the matrix T", things are ambiguous because the elements of this matrix could be Tab or Tba which are different. Similarly, matrices T-1 or TT are ambiguous. So we will define first a matrix called Tup which has these matrix elements and transpose (Tup)ab ≡ Tab [(Tup)T]ab = (Tup)ba = Tba We can do a similar definition for a down matrix Tdn in terms of the down-tilt T, (Tdn)ab ≡ Tab [(Tdn)T]ab = (Tdn)ba = Tba The orthogonality conditions were these δac = Tba Tbc // sum on first δca = Tbc Tba // sum on second and these can be written in terms of our matrix Tup and Tdn δac = (Tup)ba (Tdn)bc δca = (Tdn)bc (Tup)ba which we rewrite as δac = (Tup)Tab (Tdn)bc => 1 = TupT Tdn δca = (Tdn)Tcb (Tup)ba => 1 = TdnT Tup Assuming inverse matrices exist, we can then write 1 = TupT Tdn 1 = TdnT Tup TupT = Tdn-1 TdnT = Tup-1 (TupT)-1 = Tdn = (Tup-1)T (TdnT)-1 = Tup = (Tdn-1)T where we recall that "inverse and transpose commute". We now define two more matrices as follows: (T-1)up ≡ (Tup)-1 = (Tdn)T => (T-1)ab = Tba (T-1)dn ≡ (Tdn)-1 = (Tup)T => (T-1)ab = Tba With such a definition, we can say that "up and inverse commute". This allows us the following notation, Tup (Tup)-1 = 1 Tup (T-1)up = 1 Tab (T-1)bc = δac = (T-1)ab Tbc Tdn (Tdn)-1 = 1 Tdn (T-1)dn = 1 Tab (T-1)bc = δac = (T-1)ab Tbc where the last equality in each group results from starting with the matrices in reverse order. In the last line in each group, we see examples of "up tilt matrix multiplication" and similarly for down tilt. Appendix B. Three questions of possible interest. 1. Is Tab a "mixed tensor" as defined in Section 5? [ no ] 2. Can we raise and lower indices on Tab ? [ yes ] 3. Are the objects δab and εabc.. tensors? [ yes ] (a) Is Tab a mixed tensor? Consider this sample equation where we use the contraction tilt rule, Qac = XabYbc = XabYbc The tensors X and Y on the right are in down-tilt format, and everything is in x-space, X Y and Q. If we transform this equation to x' space using the appropriate up or down tilt T for the appropriate up or down index, we would get the same equation but everything is primed (equation is form invariant). In showing this, we make use of the orthogonality rule mentioned earlier. So we get, Q'ac = X'abY'bc = X'abY'bc and now everything is in x' space. So for each of these rank-2 tensors, Qac and Q'ac , both indices are always "in the same space". Now, the notation Tab makes it "look like" a tensor entirely in x space like Qab above, just because we have no prime on T. But this is not true. The object Tab has its first index in x' space and its second index in x space, so it is not entirely in either space. We therefore cannot transform Tab from x space to x' space the way we transformed Qab. There is no object T'ab. [ We could of course talk about (T-1)ab which would transform things from x' space back to x space. ] (b) Can we raise and lower indices on Tab ? We did not talk in the main text about the possibility of raising and lowering indices on the transformation matrix itself, Tab . It seemed like an extra complication in the middle of an already complicated flow, so we put the subject here in this appendix. First, a reminder that gab transforms in this way, g'ab = Tac Tbd gcd So given the tensor g in x space, there is some corresponding tensor g' in x' space which is going to be different from g, generally speaking. We can raise and lower indices on T provided we use the g matrix appropriate for the space of that index. Here are some examples Tab = g'abTbc raise first index of T using g' Tab = Tacgcb raise second index of T using g Equating the two expressions above we get Tacgcb = g'abTbc (*) Apply gcd on the right of both sides to get Tad = g'ab Tbc gcd (**) Alternatively, we could go back to (*) and apply g'ea from the left to get g'ea Tacgcb = Tec We have already seen this last result in Section 6 (f) above. Now we ask: is our result (**) consistent with our introduction to the T matrix forms in Sections 1 and 3 above? ( I think this has already been answered in the text, but let's do it again here). There we said: dx'a = Tab dxb Tab = dx'a = Tab dxb Tab = We claimed the first equation represented "normal physics" and introduced the second equation as the way something called a "covariant vector" transforms. But now we can in fact derive the second equation from the first equation using our new rule (**). So start with the first line of the pair above (change b to d) dx'a = Tad dxd dx'a = (g'ab Tbc gcd ) dxd = g'ab Tbc dxc Apply g'ea on the left to both sides g'ea dx'a = g'ea g'ab Tbc dxc dx'e = Tec dxc and this is in fact the second line. The upshot is that, if we want, we can raise and lower either index on Tab but we must use g' when acting on the first index, and g when acting on the second. Once again, let's simultaneously raise the first index and lower the second index g'ca Tab gbd = Tcd which gives a fast way to result (**) above. (c) Are the objects δab and εabc.. tensors? The answer depends a little on usage. Either of δij or εijk can be regarded as a "constant" bookkeeping device, or as a tensor. As a bookkeeping device, we could use either symbol in x space or x' space. But if we want to regard their indices as proper tensor indices, we need to treat them as tensors, which means that if x-space is Cartesian and x'-space is some other space, we have δ'ij = Tia Tjb δab ie g'ij = Tia Tjb gab ε'ijk = Tia Tjb Tkc εabc Looking at the first line, we realize that δ'ij is really g'ij and the notation δ'ij is a bit misleading since we expect this to vanish for different indices, but in general it does not. The second line is a bit more interesting. Notice that the object ε'ijk, whatever it is, is totally antisymmetric. We can therefore write it as ε'ijk = εijk ε'123 . Then we have, ε'ijk = εijk ε'123 = εijk T1a T2b T3c εabc = εijk det(Tab) = v εijk = εijk = εijk So the answer is yes, εijk is a tensor, and in any frame it is totally antisymmetric, but it picks up a factor of in scale as you move from the Cartesian coordinates to curvilinear coordinates. As usual, if you want to raise or lower an index on ε'ijk, you would do so with g' and not g. In our work above, we have nevered use ε'ijk . Appendix C. The Action of a Gradient on a Vector (a) Preliminary comments: We have noted that Tab is in general a function of coordinates, which means that the derivative ∂cTab ≠ 0. We could write either T(x)ab or T(x')ab to show this, but of course these are different "functions" of their argument, so let's avoid that issue by just not showing the dependence. We know that, even if we start with g = 1 in a Cartesian system, if we apply some general transformation T we will end up in the x' space with g'ab = Tac Tbc so so g'ab will depend on the coordinates, and then we will have ∂'c g'ab ≠ 0. But we might just as well assume from the start that the metric tensor gab in x-space is a function of coordinates and be completely general, so ∂c gab ≠ 0. In this case, when derivatives are combined with tensors, things can get fairly complicated, as we now show in just a simple case. This situation arises, by the way, in general relativity. (b) Gradient of a Vector: Elwin Bruno Christoffel (1829–1900). Consider application of the gradient to a vector in x' space, where we use the fact that the (regular) gradient operator ∂a transforms as a normal covariant vector, (∂'aV'b) = (Tad∂d) (TbcVc) = Tad Tbc (∂dVc) + Tad(∂d Tbc)Vc If it were not for the second term, we would say that ∂dVc was a covariant rank 2 tensor, but the second term is sitting right there, so ∂dVc is not a rank 2 tensor. Something needs to be "fixed up". If we define the following fairly complicated object, known as the covariant derivative of Vb , which has at least three different notations , aVb ≡ Vb,a ≡ Vb;a = ∂aVb – {ab,c}Vc where {ab,c} ≡ Γcab ≡ gcd [ab,d] = ½ gcd( ∂agbd + ∂bgad – ∂dgab ) // Christoffel 2nd kind where [ab,d] ≡ ½ ( ∂agbd + ∂bgad – ∂dgab ) // Christoffel 1st kind then we find, after considerable fumbling around with the algebra (see below), that V'b,a = Tad Tbc V d,c (*) and so Vb,a is a covariant rank 2 tensor. Notice that in V'b,a we will have {ab,c}' ≡ g'cd [ab,d]' [ab,c]' ≡ ½ ( ∂'ag'bc + ∂'bg'ac – ∂'cg'ab ) just to remind the reader that everything is primed on the primed side of the equation. The objects [ab,d] and {ab,c} ≡ Γcab ( in two notations) are called Christoffel "symbols" and are nothing more than what you see. These particular symbols are all symmetric in the ab indices. If g = constant, all these symbols vanish and life is easy. M&M refer to Vb,a as "the comma notation", while others use aVb with the understanding that the a symbol is not just the normal ∂a = ∂/∂xa derivative. The "semicolon notation" is also used. The covariant derivative of a contravariant vector Vb is a mixed rank 2 tensor and is given by aVb ≡ Vb,a ≡ ∂aVb + {ab,c}Vc so the only difference in the form of the equation is one sign. (c) Some technical Facts. Verifying equation (*) above requires a lot of work. There are various intermediate "facts" which are helpful to know about, and which I record here. Details of the derivation are in a "support" doc. Fact 1 : (∂a Tbc) = ∂a = ≡ Sbac // symmetric on a↔c ! The invented symbol Sbac shown here is symmetric on the indices a and c, which is not so obvious when you just stare at (∂a Tbc) . So it serves as a little crutch. Fact 2: (∂s Tce) = – Tae Tcb Sasb Notice that only the up-tilt T's appear here, and our S object from Fact 1. Fact 2 is derived by setting the derivative of the two orthogonality equations to 0, which starts one off as follows: 0 = ∂s(Tab Tcb) = (∂sTab) Tcb + Tab (∂s Tcb) 0 = ∂s(Tab Tac) = (∂sTab) Tac + Tab (∂s Tac) Fact 3: (∂s gdc) = – gdb gca(∂s gab) This one can be derived by differentiating gab gac = bc which is to say ∂s (gab gac) = 0. Fact 4: Qab ≡ Tad(∂d Tbc) = – Tad Tgc Tbf Sgdf = – Tgc Tad Tbf Sgdf symmetric in ab The object Qab defined as the left expression does not look symmetric in a and b, but you can use Fact 2 to get the subsequent forms and then the last form reveals the symmetry. Since Sgdf is symmetric on d and f, the expression shown is symmetric on a and b.