Phil Lucht Math & Physics Archive
Home / Math and Physics Files / Math / Wedge World / tensor wedge doc / tensor obs

tensor product definition section

DOCX · 98.3 KB
Open DOCX file

Expository section from Phil's Wedge World tensor documents, marked closed and not to be edited. Its outline covers the tensor product as a quotient space, category theory, outer products, tensor algebra and Kronecker products. The visible text builds F(VxW)/N from the Cartesian product, derives the bilinearity rules, shows non-commutativity, verifies the vector space axioms, finds a basis and dimension nn', and generalizes to several spaces.

AI-written summary; may contain errors.

Extracted text (machine-read; may contain errors)
this doc is closed do not edit 1. The Tensor Product 1 1.1 The Tensor Product as a Quotient Space 1 1.2 The Tensor Product in Category Theory 5 1.3 Outer Products 7 1.4 Tensor products of the form VV...V : The Tensor Algebra 9 1.5 Kronecker Products 12 1. The Tensor Product There are two theoretical paths leading to the tensor product concept which we briefly summarize here in a non-rigorous manner, then we extend the concept to outer products and Kronecker products. Nomenclature: Tensor Product vs Direct Product. The tensor product described in this section sometimes goes by other names. In quantum mechanics, a system of two particles might be in a quantum state |ψ1> |ψ2> which is an element of a tensor product space V1V2 (as we shall describe below). Some quantum authors refer to this tensor product as a direct product (e.g. Shankar) while others call it a tensor product (e.g. Messiah). It happens that in quantum theory states like |ψ1> reside in a vector space which is also a Hilbert space. Similarly, when a quantum system has a symmetry, such as rotational invariance (e.g., an isolated atom), the quantum states can be classified into certain vector spaces associated with the matrix representations of the symmetry group, and the tensor products of these spaces are usually called direct products. For example, the rotation group has matrix representations called 1 and (1/2), and one says 1 (1/2) to indicate the "direct product" of these two spaces. It happens that the matrix for 1 (1/2) can be brought into diagonal block form by a similarity transformation, where the blocks are (1/2) and (3/2), and so one finds that 1 (1/2) = (1/2) (3/2). Sometimes the tensor product is called a tensor direct product which seems associated with the outer product extension of the tensor product noted below. Sometimes the raw Cartesian product (see below) is called a direct product, but usually there is some additional structure involved. Generally, the term direct product seems most suitable for the direct product of groups, rings, modules and related objects, whereas in the current document we are discussing the tensor product of vector spaces and of the tensors contained within those spaces. Category Theory mentioned below attempts to put all these product into a uniform framework. 1.1 The Tensor Product as a Quotient Space It does seem odd that one might think of a product VW in terms of a quotient. We shall outline how this path goes in a set of steps. 1. Cartesian Product. Start with the inert Cartesian product set VxW with elements (v,w), where in our application the sets V and W are vector spaces. This set VxW is "inert" in the sense that one has no instructions for what can be done with its elements. 2. Space F(VxW). We now endow VxW with an addition operator + and a scalar multiplication operator (indicated by juxtaposition) allowing us to form linear combinations of elements with scalar coefficients. Let's define F(VxW) to be a space which contains all such linear combinations. A typical element of this F(VxW) space might be 3(v1,w3) - 2.1(v2,w5). Of course (v1,w3) also lies in F(VxW) and one might call this a pure element, whereas 3(v1,w3) - 2.1(v2,w5) is a mixed element. Because the sum of two linear combinations is again a linear combination of the same form, the space F(VxW) is closed under addition. 3. Field K. We usually assume (as above) that the scalars are in the field R of real numbers, but to be more general one can assume the scalars are elements of some arbitrary field traditionally called K. In addition to the reals R, there are various fields having an infinite number of elements (like rational or complex numbers), and there are various fields of having a finite number of elements (the Galois Fields). Footnote: Sometimes the space F(VxW) is described as a "free vector space" which is a set of functions f such that f: VxW → K. Usually such spaces are defined over a discrete set S, and it is not clear how this works when the set S is continuous, this being the case for S = VxW. Moreover, the functions f mapping to K cannot be identified with our linear combinations since for example (v2,v5) is not an element of K. We therefore refrain from giving F(VxW) this moniker and the reader should regard F(VxW) only as we have defined it above. 4. Equivalence Relations and Classes. Now define the following set of "equivalence relations" (v1+v2, w) ~ (v1,w) + (v2,w) for all v1,v2 ϵ V and all w ϵ W (v, w1+w2) ~ (v,w1) + (v,w2) for all v ϵ V and all w1, w2 ϵ W α(v,w) ~ (αv,w) for all v ϵ V and all w ϵ W and all α ϵ K α(v,w) ~ (v,αw) for all v ϵ V and all w ϵ W and all α ϵ K where ~ means "is equivalent to". Rewrite these relations as (v1+v2, w) – (v1,w) – (v2,w) ~ 0 (v, w1+w2) – (v,w1) – (v,w2) ~ 0 α(v,w) – (αv,w) ~ 0 α(v,w) – (v,αw) ~ 0 . We are declaring here that lots of linear combinations are equivalent to 0. The reason we do this is to make our tensor product space (to be defined below) have "nice properties" (i.e., it is then a vector space). These functions taken together define an "equivalence class" which is equivalent to 0. Call this class N (for null). 5. The Quotient F(VxW)/N. There then exists a space which we shall call F(VxW)/N, or F(VxW) "mod" N. This is a standard structure in equivalence class theory where one takes the quotient of one space S divided by another space of equivalent items in space S, often written S/~. The upshot is that the elements of the new quotient space F(VxW)/N consist of all linear combinations of F(VxW) except that any linear combination which has one of the four forms shown above is filtered out ("modded out") by setting it equal to 0. Example: 3(w3,v4) + (v2, w1+w2) – (v2,w1) – (v2,w2) = an element of F(VxW) 3(w3,v4) = the corresponding element of F(VxW)/N. 6. The Space VW. We now give this space F(VxW)/N a new name: F(VxW)/N = VW = the tensor product space of V and W . The elements of VW are linear combinations of elements called vw instead of (v,w) as a reminder that the equivalence class N must be respected. Whereas the comma in (v,w) was a mere separation operator, the in vw is regarded as a new "tensor product multiplication operator" with the properties listed below which, in effect, implement the equivalence relations stated above. 7. Practical Summary. The end result of all this song and dance is the following: The tensor product space VW is the set of all linear combinations of elements (v,w) of the Cartesian product set VxW, written as vw, where the following rules are declared by fiat: (v1+v2) w = (v1w) + (v2w) for all v1,v2 ϵ V and all w ϵ W v (w1+w2) = (vw1) + (vw2) for all v ϵ V and all w1, w2 ϵ W α(vw) = (αv)w for all v ϵ V and all w ϵ W and all α ϵ K α(vw) = v(αw) for all v ϵ V and all w ϵ W and all α ϵ K The first two rules state that multiplication is distributive over addition (from right and left), while the last two rules state the scalars work in the expected manner. If these rules were declared for a function f(v,w), they would appear as f(v1+v2,w) = f(v1,w) + f(v2,w) f(v,w1+w2) = f(v,w1) + f(v,w1) α f(v,w) = f(αv,w) α f(v,w) = f(v,αw) Such a function would then be described as being bilinear because it is linear separately in each argument. One can then regard the rules shown above for as expressing bilinearity for the tensor product space VW. 8. vw does not commute. Whereas the + operation within VW is commutative, it should be clear that the operation is not commutative. If v ϵ V and w ϵ W, then vw ϵ VW whereas wv is an element of a completely different space which is WV. Even if W = V, one has vv' ≠ v'v if v≠v'. The fact goes back to the original Cartesian product set VxV where one has (v,v') ≠ (v',v) if v≠v' because (v,v') is an ordered tuplet, not a set {v,v'}. If V = W = R, one would not identify the point (x,y) with the point (y,x) in RxR = R2 if x ≠ y. Another word for commutative is abelian. 9. VW is a vector space. The space VW is a vector space whose vectors are linear combinations of vw. We shall now verify this to be the case. We already know VW is closed under addition since F(VxW) had this property. The inverse of vw is (-1)(vw). Addition is commutative and associative. Any element of the form 0w or v0 can be taken as the identity for addition (the "zero") since, for example, 0w = (v - v) w = (vw) + ((-v)w) = (vw) - (vw) = 0 (0 in the space VW) . There is a scalar multiplicative identity since all fields K have an identity "1": 1(vw) = (vw). "Vector multiplication" is distributive over scalar addition (here the "vector" is vw), (α + β)( vw) = [(α+β)v]w = [αv+βv]w = (αv)w + (βv)w = α(vw) + β(vw). Multiplication by a scalar is distributive over "vector addition" : α (v1w2 + v3w4) = α (v1w2) + α (v3w4) This property we more or less add by fiat to the earlier properties. It is the only reasonable way to do things since elements of VW are linear combinations of pure elements of the form vw. 10. Basis of VW and general elements of VW. In the above verification that VW is a vector space, we used only pure vectors of VW, but general vectors of VW are linear combinations of the pure vectors so we really should rehash the above for general vectors. To do this, we first note that, since V and W are vector spaces, each has a basis, and we call these bases {ei} for V and {e'i} for W. It is not hard to show that the set of elements of the form eje'j forms a basis for VW, so a general vector in VW can be expressed as u = Σij uij (eie'j) uij ϵ field K, "coefficients" . The inverse element -u is pretty obvious. Addition u + u' is commutative and u + u' + u" is associative. The zero element is the same. Vector multiplication is still distributive over scalar addition, (α + β)u = (α + β)[ Σij uij (eie'j)] = Σij uij [ (α + β) (eie'j)] = Σij uij [ α (eie'j) + β(eie'j)] = α [ Σij uij (eie'j)] + β [ Σij uij (eie'j)] = α u + β u . In this manner, all the required properties of a vector space can be verified for general elements of VW. 11. Vector vs Tensor. Since VW is a vector space, it is proper to refer to its elements vw (or linear combinations of same) as "vectors". On the other hand, we shall refer to vw as a "tensor" in the tensor product space VW. In particular, it is a "rank-2 tensor" composed from v and w which are vectors in their respective vector spaces V and W. The word vector must be evaluated in its context. The notion of tensors is developed more in Section 1.4 below. 12. Dimension of VW. As noted above, the basis of the vector space VW consists of elements of the form eie'j . If the dimensions of V and W are n and n', then i takes n values, j takes n' values, and the dimension of the vector space VW is nn', the product of the separate vector space dimensions. 13. Generalization. The above development is easily generalized to the tensor product of any finite number of vector spaces. One first defines F(V,W,....Z) as linear combinations of elements of the Cartesian product space VxWx..xZ , which elements have the form (v,x,...z). One then defines a large set of equivalence relations analogous to those described above. One ends up with a large set of linear combinations which are all equivalent to 0, and this defines the equivalence class N. One then creates F(V,W,....Z)/N as the space of linear combinations where any pieces which are equivalent to 0 are filtered out. One then defines F(VxWx...xZ)/N = VW...Z = the tensor product of spaces V and W and... and Z. The tensor product space VW...Z is the set of all linear combinations of elements (v,w,...z) of the Cartesian product space VxWx...xZ, written as vw...z, where the following rules are declared by fiat: (v1+v2)w .... z = v1w .... z + v2w .... z v(w1+w2) .... z = vw1 .... z + vw2 .... z and so on for all positions in the product. In addition, α(vw...z) = (αv)w...z α(vw...z) =v(αw)...z and so on for all positions in the product. When these rules are written for a function f(v,w,....z) one has, f(v1+v2,w,...z) = f(v1,w,...z) + f(v2,w,...z) f(v,w1+w2,...z) = f(v,w1,...z) + f(v,w2,...z) etc α f(v,w,...z) = f(αv,w...z) = f(v,αw,...) = etc. If there are k factors in the tensor product VW...Z, then the function f has k arguments, and a function obeying all of the above rules is said to be k-multilinear. For k = 2 we have bilinear, for k = 3 we have trilinear, and so on. One can mix in the scalar rule by saying for example f(αv1+βv2,w,...z) = αf(v1,w,...z) + βf(v2,w,...z) and similarly for all positions but we have kept the scalar rules separate. Either statement of the rules is equivalent. We can then regard the set of rules shown above as describing k-multilinearity for the tensor product space VW...Z. Written in the second form we would say (αv1+βv2)w .... z = α (v1w .... z) + β (v2w .... z) and similarly for all positions. An alternate approach to developing the tensor product of three or more vector spaces is to inductively build up by grouping things. For example VWX = (VW) X = the tensor product of two vector spaces, one of which is VW VWXY = (VWX)Y = the tensor product of two vector spaces, one of which is VWX The results are the same with either approach. 1.2 The Tensor Product in Category Theory Category theory is an attempt to abstract the essence of algebraic structures which apply generally to objects like vector spaces, sets, rings, groups, modules and so on. One encounters certain category diagrams which must allow for flow through the diagram in all possible ways (the diagram must "commute"). A diagram consists of certain objects which are connected by arrows known as morphisms. For our application, these arrows are function mappings between spaces, and two sequential arrows in a path represent function composition in the sense f o g. At a higher level, if the objects in the diagram are themselves categories, the morphism arrows are called functors, a concept we happily won't need. Category theory is a relatively recent addition to the house of many mansions. With precursor work done by Emily Noether (whose work shows up in a lot of places), category theory was developed in the early 1940's by Saunders Mac Lane (and others) who then summarized the theory in a text Algebra (1967) with coauthor Garrett Birkhoff. These same authors wrote the classic textbook A Survey of Modern Algebra (1941/1997) which is known to many students as "Birkhoff and MacLane". We give here just an outline of this rather slippery tensor product development. It seems more of a fitting of our conclusion of Section 1.1 into category theory. The reader interested in more detail can look at Chapter 14 "Tensor Products" of Roman's text Advanced Linear Algebra (2007). We start with the following triangle diagram (an example of a category diagram), In this diagram VxW is the Cartesian product of two vector spaces V and W, exactly as used in Section 1.1 above. Elements of VxW are (v,w). There are two mappings f: VxW → X and g: VxW → Y where X and Y are for the moment just spaces. They in turn are linked by a mapping traditionally called τ, so τ : X → Y. One says that "g can be factored through f". The functions f and g are declared bilinear from the get-go. This is analogous to our declared equivalence relations in the approach of Section 1.1. The set of all bilinear mappings f: VxW→X is called homF(V,W; X) where F is the field of scalars. The letters hom stand for homomorphism ("same shape") which is a structure-preserving map. Linear maps (like τ discussed below) preserve vector space structure. The triangle diagram must commute, so we must have g = τ o f (function composition). The space X is our candidate space for the tensor product VW space. We need to construct the function τ. To do so, we use the fact that the diagram commutes to evaluate τ at the pure points vw, τ(vw) = g(v,w). Now we "extend" τ so it applies to linear combinations of vw elements by declaring that τ (α vw + β v'w' ) = α τ(vw) + β τ(v'w') = α g(v,w) + βg(v',w') so τ is now a linear function τ: X→Y. It maps every element of X into an element of Y, and it is unique by its construction. Once we have τ being a unique linear mapping, the "pair" (X, f:VxW→X) becomes a "universal pair". The idea here is that any alternate "pair" like (Y, g:VxW→Y) is equivalent to (X, f:VxW→X) up to the isomorphism implied by τ. In this sense, then, the mapping f:VxW → VW is essentially unique -- it is "universal for bilinearity" -- so the tensor product mapping is well-defined. Function τ is called a mediating morphism, f is called the tensor map, and the elements of VW are tensors. Our "rules" of Section 1.1 for operator now derive from the fact that g is a bilinear function: (v1+v2) w = g(v1+v2,w) = g(v1,w) + g(v2,w) = (v1w) + (v2w) v (w1+w2) = g(v,w1+w2) = g(v,w1) + g(v,w2) = (vw1) + (vw2) α(vw) = α g(v,w) = g(αv,w) = (αv)w α(vw) = α g(v,w) = g(v,αw) = v(αw) We then end up with the same space VW and rules as in the previous development, and we have extra assurance that VW is a unique and well-defined object. The above scenario directly generalizes to the tensor product of k vector spaces with the following corresponding category diagram, Lang (Algebra) for example shows the equivalent of this diagram on page 602 of his Chapter 16 (The Tensor Product). 1.3 Outer Products By either development above, we have a definition of the tensor product VW with pure elements of the form vw, and with general elements of the form F = Σij Fij(eie'j), along with the bilinear rules noted earlier, (v1+v2) w = (v1w) + (v2w) for all v1,v2 ϵ V and all w ϵ W v (w1+w2) = (vw1) + (vw2) for all v ϵ V and all w1, w2 ϵ W α(vw) = (αv)w = v(αw) for all v ϵ V and all w ϵ W and all α ϵ K . The essence of the tensor product is this bilinearity, and there is no requirement to describe the objects vw in more detail. In the formal sense we are done and fini. However, for "engineering purposes", it is useful to add more structure to the tensor product by defining "tensor components" such as [vw]ij ≡ viwj where vi and wj are components of vector v in V and w in W. These vi and wj can be the elements of any field K and the juxtaposition of viwj implies multiplication in that field. but we have in mind that K = R, the real numbers. The key point: because the function viwj is manifestly bilinear, this extra specification does not conflict with any of the earlier tensor product "rules". For example we can evaluate, [(v1+v2) w]ij = (v1+v2)iwj (v1w)ij + (v2w)ij = (v1)i wj + (v2)i wj The first rules says the left sides of these two equations must be equal, but we can see that the right sides are also equal, so things are consistent. When this component level structure is glommed onto the raw tensor product, we end up with something called an outer product, though it is indicated by the same notations VW and vw. Thus an outer product is a tensor product which has been "extended" so we can talk about components of the various tensors. Most readers probably think of tensors a priori as having "components", but here we have made a thin distinction between tensors "in the large" and tensors with components which can be manipulated. Consider then the following most-general expansion of a tensor T of VW, T = Σij Fij eie'j . Suppose the basis vectors ei are simple "unit vectors" so (ei)k = δi,k, and similarly for the e'i of W, (e'i)k = δi,k ( but remember V has n axes and W has n' axes). Then we can take the "ab-component" of the above tensor expansion as follows : Tab = [ Σij Fij eie'j]ab = Σij Fij [eie'j]ab = Σij Fij (ei)a(e'j)b = Σij Fij δi,a δj,b = Fab . Thus we see that the expansion coefficients Fij in this simple basis are exactly the tensor components Tij. The same thing happens in the V world: v = Σiaiei vk = Σiai(ei)k = Σi aiδi,j = ak . _________________________________________________________________________________ Exercise: Show how Fij and Tij are related for a general basis eie'j. It will be shown in Section ** that the bases ei and e'i have dual bases qi and q'i such that, in the matrix/vector notation of ***, qiT ej = δi,j q'iT e'j = δi,j which can also be written eiT qj = δi,j e'iT q'j = δi,j . For the simple basis (ei)k = δi,k one has qi = ei, but for other bases, the qi are linear combinations of the ei. Let is reconsider then in a completely general basis eie'j, T = Σij Fij (eie'j) Tab = Σij Fij (eie'j)ab = Σij Fij(ei)a(e'j)b . Multiply both sides by (qmq'n)ab = (qm)a(q'n)b and sum on a and b, Σab (qm)a(q'n)bTab = Σab Σij Fij (qm)a(q'n)b(ei)a(e'j)b = Σij Fij [Σa(qm)a(ei)a] [Σb(q'n)b(e'j)b] = Σij Fij δm,i δn,j = Fmn so Fij = Σab (qi)a(q'j)bTab ** Meanwhile, a basis is complete with respect to its dual basis (as we show just below) so we also know that Σi(qi)a(ei)b = δa,b Σi(q'i)a(e'i)b = δa,b . We can then multiply both sides of ** by (ei)m(e'j)n and sum on i and j to get Σij(ei)m(e'j)nFij = Σij(ei)m(e'j)n Σab (qi)a(q'j)bTab = Σab Tab [Σi (qi)a(ei)m ][Σj(q'j)b(e'j)n] = Σab Tab δa,m δb,n = Tmn Our conclusions just obtained can be written as Fij = Σab (qi)a(q'j)bTab // sum is on component indices Tij = Σab(ea)i(e'b)jFab // sum is on basis vector labels Thus the coefficients Fij are linear combinations of the Tij, and vice versa, with coefficients as shown. We can write the first equation the matrix/vector notation of *** as Fij = (qi)T T (q'j) For comparison, we found earlier that Tij = (ei)T T (e'j) . In each case, one can think of the tensor T as a basis-independent operator which we sandwich between different sets of vectors to create the Fij and Tij. In the case W = V, one would write the above equations in quantum mechanics notation as Fij = <qi | T | qj > = matrix elements of operator T in the {qi} basis which is dual to general {ei} Tij = <ei | T | ej > = matrix elements of operator T in the simple {ei} basis Completeness relation. Write v = Σiviei. Apply qjT from the left to get qjTv = qjT(Σiviei) = ΣiviqjTei = Σiviδj,i = vj Then we have v = Σiviei = Σi[qiTv]ei so va = Σi[qiTv](ei)a or va = Σi[Σb(qi)bvb](ei)a = Σb vb [ Σi(qi)b(ei)a] . For this equation to be valid for any component of any vector, we must have [ Σi(qi)b(ei)a] = δb,a Some authors write this completeness relation as ΣiqieiT = 1 with the meaning as shown above. _____________________________________________________________________________ What we find is that by adding information to the tensor product to create an outer product, we create machinery which is useful in component manipulations of tensors. This is especially useful when we have V = Rn and W = Rn', and even more useful when V = W = Rn. Usually the term "rank-2 tensor" assumes that V = W. The idea of an outer product can easily be extended to the tensor product of more than two spaces. Here are upgrades of some equations above for the case k = 3 where we have VWX : [vwx]ijk ≡ viwjxk // manifestly trilinear! T = Σijk Fijk (eie'je"j) . // most general element of VWX (e"j are a basis for X). Fijk = Σabc (qi)a(q'j)b(q"k)cTabc The generalization to an outer product of k vector spaces is then the following: [vwx....]ijk ≡ viwjxk ..... // manifestly k-multilinear T = Σijk... Fijk... (eie'je"j...) . // most general element of VWX .... Fijk... = Σabc... (qi)a(q'j)b(q"k)c....Tabc... Below we shall extend the idea of an outer product to apply to tensors other than vectors. 1.4 Tensor products of the form VV...V : The Tensor Algebra We have tried to keep things general up this point by using VW...Z where all the vector spaces could be different, but in this section we assume they are all the same, and this is our main interest. There is then only one set of basis functions {ei} to worry about, the basis for V. We now introduce the compact notation: Vk ≡ VV....V // k copies, fancy notation Πi=1k V In Section 1.3 we defined the outer produce as a specialization of the tensor product where [abc....]ijk... = aibjck...... which is a manifestly k-multilinear function. Here a,b,c... are all vectors in V. We would like to generalize the notation so it can act on objects other than vectors. This is fairly easy to do. We start by defining: Aij ≡ [ab]ij = aibj rank-2 tensor We refer to Aij as the components of a rank-2 tensor A = ab, whereas ai and bi are components of a rank-1 tensors a and b. The word "tensor" has a weak and a strong meaning as discussed in ***. In the weak meaning, a rank-2 tensor is something that has components with two indices like Aij. In the strong meaning, a rank-2 tensor is a set of components Aij which transform in a certain manner with respect to some transformation. This is the meaning one works with when the ei are the tangent base vectors of a transformation, and the ei are the reciprocal base vectors, as already discussed. In either sense of the word tensor, the statement Aij = aibj defines a rank-2 tensor as the outer product of two rank-1 tensors, and this is a profitable way to construct higher rank tensors from lower rank ones. The outer product form also specifies exactly how a rank-2 tensor should "transform" in terms of how rank-1 vectors "transform" under a transformation. For example under a rotation a vector goes as v'a = ΣiRaivi and then the transformation of the rank-2 tensor is A'ab = ΣijRaiRbjAij . A prime here indicates the vector or tensor that has been rotated and thus has different components. The term "rank-2 tensor" also indicates the total object A which has components Aij, just as a "vector" is a total object v which has components vi. Some authors replace the word "rank" with "order" so then A is a tensor of order 2, or perhaps A is a "2-tensor" and the words rank or order are avoided altogether. We shall use the word "rank" even though it has other unrelated mathematical meanings (e.g. rank of a matrix). It should be understood that not all rank-2 tensors can be written as outer products of vectors. For example, in the general expansion T = Σij Tij eiej in the simple basis noted above we have T being the total rank-2 tensor, and Tij being its components. But in general, we don't have Tij = aibj. The key point is that Tij "transforms" and "behaves" the same way as aibj. We could construct an explicit Tij = aibj if we wanted by saying a b = (Σiaiei) (Σjbjej) = Σij (aibj) eiej. So, let's first define two outer product rank-2 tensors A = ab Aij = [ab]ij = aibj B = cd Bij = [cd]ij = cidj We can then define the tensor product of A and B in a fairly obvious manner, AB ≡ (ab)(cd) = abcd . // associative Then we take components in the following way, [AB]ijkl = [ abcd]ijkl = aibjckdl = AijBkl The above lines shows that AB is in fact a rank-4 tensor constructed by taking the "outer product" of two rank-2 tensors (or the outer product of four rank-1 tensors). Later we shall have use for an object defined in this strange manner [AB]ik,jl ≡ [AB]ijkl = AijBkl and we just mention it here in passing. Note that the indices are shuffled. Using the same method as above, we could construct a rank-6 tensor from the tensor product of three rank-2 tensors, [ABC]abcdef = AabBcdCef or from the tensor product of two rank-3 tensors [AB]abcdef = AabcBdef . At a slightly lower level, we could construct a rank-3 tensor from a rank-1 tensor and a rank-2 tensor, another form of the outer product, [v B]abc = va Bbc . In general, one can take the tensor product of any set of tensors to create a new tensor whose rank is the sum of the ranks of the tensors that were combined by the symbol. If α,β,γ... are arbitrary tensors, having multiindices I,J,K (for example I = {i1,i2,i3} if α is rank-3), we could write a general formula for the components of the tensor product of any number of pure tensor objects in this manner, [ α β γ ....]IJK... = αI βJ γK...... (*) The tensor here is α β γ .... and the equation specifies its components. The Direct Sum of Vector Spaces The direct sum of two or more vector spaces is a very simple concept. Consider G ≡ V2V3. The vectors in this space have 5 components and can be visualized by just stacking vectors into tall column vectors. For example, v = ϵ V2 w = ϵ V3 = ϵ V2V3 . One could write the above as vw ϵ V2V3. Since α = , one has α(vw) = (αv)(αw). And to the extent that a vector's position within the larger vector is immaterial, one has vw = wv . Regardless, one sees that each piece of the direct sum vector resides in its own private region within the tall column vector. Whereas dim(AB) = dimA * dimB, one has dim(AB) = dimA + dimB. The Tensor Algebra Suppose we define a very large vector space as the following direct sum of vector spaces, T = V0 V V2 V3 ....... on forever // = Σi=0∞ Vi The space V0 represents the space of scalars. The elements of this space T are then tensors of any rank. A vector would fit into the V part, a rank-2 tensor would fit into the V2 part, and so on. One could in fact have a linear combination of a vector and a rank-2 tensor in this space T. Since T contains arbitrary linear combinations of tensors of different ranks, T is called a "graded algebra" where the ranks of the pieces are the grades. The "algebra" part is that we have a set elements with rules + and for adding and multiplying elements. The set of elements of T is closed under each of these operations. This last fact should be clear to the reader. For example, one can think of the product of any number of tensors as [**...][**...][**...] ..... = ***** .... ϵ Vk (if k factors) ϵ T Then the most general elements of the space T are linear combinations (with coefficients in field K) of the tensors shown in (*). Thus huge vector space with all its many elements is known as the tensor algebra over field K", and again, we usually use K = R = reals. Of course vectors in the space T defined above have an infinite number of components. Usually only a finite number of these components are non-zero. In the context of this large space T, we have two meanings for "mixed" tensors which we don't want to confuse together. A pure rank-2 tensor has the form ab. A mixed rank-2 tensor is ΣijFijei ej . But a "mixed tensor" in the space T could be any linear combination of tensors of different rank, such as 5.3 + 2v + ΣijFijei ej - π abc Suppose t is such a mixed tensor composed of a linear combination of tensors whose ranks range from s to r. And suppose t' has ranks ranging from s' to r'. Then tt' can contain tensors whose ranks range from |s-s'| to r+r' 1.5 Kronecker Products The subject here is the tensor product of two linear operators, but as will be seen, it boils down combining two rank-2 tensors to make a rank-4 tensor. The reader can regard this section as an exercise in using the tensor product machinery, and the Kronecker product just arises along the way. Let V and X be vector spaces of dimension n and m. Basis(V) = ei Basis(X) = ei Let W and Y be vector spaces of dimension n' and m'. Basis(W) = e'i Basis(Y) = e'i Consider linear operators S and T such that, x = Sv = a vector in X S: V→X xi = Σa=1n Siava i = 1,2..m y = Tw = a vector in Y T:W→Y yj = Σb=1n'Tjbwb j = 1,2..m' The linear operator S is represented by matrix Sia which has m rows and n columns (m x n). The linear operator T is represented by matrix Tjb which has m' rows and n' columns (m' x n'). We want to create a meaning for ST which is the tensor product of these two operators S and T. A candidate definition for this meaning is the following, (ST)(vw) = (xy) = (Sv)(Tw) ST : VW → XY . Consider the following processing steps, (ST)([αv1 + βv2]w) = (S[αv1 + βv2])(Tw) // definition of action of (ST) = (α Sv1+ βSv2) (Tw) // S:V→X is linear = α (Sv1)(Tw) + β(Sv2)(Tw) // using the first rule for vectors = α (ST)(v1w) + β (ST)(v2w) . // definition of action of (ST) This shows that (ST)(vw) is linear in v. A similar argument shows it is also linear in w. Thus, the operator (ST) as defined above is a bilinear operator on VW, and we confirm the essential characteristic of the tensor product, which is its bilinearity. We accept the candidate. ____________________________________________________________________ Exercise: Compute the action of (ST) on a general element of VW . Applying (ST) to a general element of VW we get (ST)[ ΣijFij eie'j] = ΣijFij (ST)( eie'j) = ΣijFij (Sei)(Te'j) . The action of S on a vector v (and T on w) can be written as (Sv) = Σa[Sv]aea = Σa(ΣbSabvb)ea = Σab(Sabvb)ea (Tw) = Σc[Tw]ce'c = Σc(ΣdTcdwd)e'c = Σcd(Tcdwd)e'c . Therefore (Sei) = ΣabSab(ei)b ea (Te'j) = ΣcdTcd(e'j)d e'c . Then (Sei)(Te'j) = [ ΣabSab(ei)b ea] [ ΣcdTcd(e'j)d e'c] = Σabcd Sab(ei)bTcd(e'j)d (eae'c) and so (ST)[ ΣijFij eie'j] = ΣijFij (Sei)(Te'j) = Σijabcd FijSab(ei)bTcd(e'j)d (eae'c) = Σac { Σijbd FijSab(ei)bTcd(e'j)d } (eae'c) = Σac Gac (eae'c) where Gac = Σijbd FijSab(ei)bTcd(e'j)d . In the special case that ei and e'j are the standard unit vector bases for V and W, the result simplifies, Gac = Σijbd FijSabδi,bTcd δj,d = Σij FijSaiTcj = Σij SaiFijTTjc = (SFTT)ac or G = SFTT . // G(m x m') = S(m x n) F(n x n') TT (n' x m'), so matrices "conform" This shows explicitly how bilinear operator (ST) acts on a general element of VW to produce an element of the space XY which has basis eie'j. _______________________________________________________________ It is useful now to consider the component analysis of the action of ST on a pure element of VW in the sense of outer products. Then (xy) = (ST)(vw) = (Sv)(Tw) so (xy)ii' = [(ST)(vw)]ii' = [(Sv)(Tw)]ii' . (**) The right side of (**) is easily processed as above, [(Sv)(Tw)]ii' = (Sv)i(Tw)i' = (Σj Sijvj)(Σj'Ti'j'wj') (*) = Σjj' SijTi'j' vjwj' = Σjj' SijTi'j' (vw)jj' . (***) It is helpful to visualize the object in the middle of (**) in this manner, [(ST)(vw)]ii' = Σjj' (ST)ii',jj' (vw)jj' (****) so then (**) becomes (xy)ii' = Σjj' (ST)ii',jj' (vw)jj' as if we were multiplying a vector (vw) by a matrix (ST), but the usual summation index is replaced by two summation indices j and j'. In a multiindex notation one might write the above as (xy)I = [(ST)(vw)]I = ΣJ (ST)I,J (vw)J I = {i,i'} J = {j,j'} . Comparing (***) and (****) we find that, (ST)ii',jj' = SijTi'j' Note carefully how the indices are arranged: S gets the firsts, T gets the seconds. This same equation appeared earlier in Section 1.4. The object (ST)ii',jj' is a rank-4 tensor, since it is the outer product of two rank-2 tensors Sij and Ti'j', and as such it has four indices. As shown earlier, normally one would write (ST)iji'j' with no comma and with the indices in the same order as those in SijTi'j'. The alternate comma notation allows the quasi-matrix multiplication point of view shown above. Is there some way to write ST as a standard matrix with two indices instead of four? Go back to our equation (xy)ii' = Σjj' (ST)ii',jj' (vw)jj' or (xiyi') = Σjj' (ST)ii',jj' (vjwj') (ST)ii',jj' = (SijTi'j') . We want to write this somehow in a form q'r = Σs Mrs qs . For illustration purposes, assume n = 2 and n' = 3. Then write the components (vjwj') as a single column vector in this obvious manner, where the w component index moves fastest, = = q with components qs where s = 1,2....n*n' . If vjwj' → qs, one can compute s from j,j' as follows: ( here 3 = n' = dim(W) for this special case ) s = (j-1)3 + j' (s-1) = (j-1)3 + (j'-1) = (j-1) + int() = j-1 and rem () = j'-1 . Thus for general n' we can compute j and j' from s in this way (integer part and remainder) j = 1+int( ) j' = 1+rem( ) . s = 1,2....n*n' One can similarly consider xiyi'→ q'r where the column vector q' has m*m' components. The rules here are i = 1+int( ) i' = 1+rem( ) . r = 1,2...m*m' Therefore, the desired Mrs is given by Mrs = (ST)ii',jj' = SijTi'j' where i = 1+int( ) j = 1+int( ) s = 1,2....n*n' i' = 1+rem( ) j' = 1+rem( ) r = 1,2...m*m' . It is a bit tedious to compute and display one of these M matrices by hand, so we let Maple do it for us. For this example we use S = m x n = 2 x 3 rows = m*m' = 6 T = m' x n' = 3 x 4 cols = n*n' = 12 Symbolically we can write M = (ST), with the meaning shown above. Matrices M of this general type are known as Kronecker products. Staring at the above matrix, one can see that the T submatrix is repeated many times, and one can write this matrix in a shorthand notation as M = where T = . This provides an easy way to manually construct such matrices. This construction is explained if we look back at the M matrix definition, Mrs = (ST)ii',jj' = SijTi'j' where i = 1+int( ) j = 1+int( ) s = 1,2....n*n' i' = 1+rem( ) j' = 1+rem( ) r = 1,2...m*m' . The indices i,j on S select a rectangular subregion of the M matrix due to their integer part definitions. Then within each subregion the i'j' indices run through their full ranges so a copy of matrix T appears in that subregion, multiplied by the Sij for that subregion. One is commonly interested in the case where S: V→V S = n x n matrix T: W→W T = n' x n' matrix With n = m = 2 and n' = m' = 2 the above code generates this matrix M, which can be compared with a result quoted on the wiki tensor product page.