Notes on Tensor Products
DOCX · 2.3 MB
Open DOCX file
Phil's working notes dated 9.16.15, last updated 9.19.15, with questionable items flagged in red. They first review his tensor document's Appendix E (direct product notation, tensor expansions, dyadics) and conclude it really uses the Kronecker/outer product. They then examine free vector spaces (wiki, Drexel and others), the wiki tensor product page, Lang, Ganatra, and Roman's universal pair and bilinear map treatment.
AI-written summary; may contain errors.
Extracted text (machine-read; may contain errors)
Notes on Tensor Products PhL 9.16.15
last update 9.19.15
I will highlight questionable items in red and add comments in red.
1. Tensor Doc use of tensor product in Appendix E 1
E.1 Direct Product Notation 1
E.2 Tensor Expansions and Bases 3
E.4 Dyadic Products 4
2. Connection between tensor doc and my current tensor thread. 4
4. A Free Vector Space 5
4.1 Free Vector Space according to the wiki tensor product page 5
4.2 Free Vector Space according to Drexel 6
4.3 Free Vector Space according to Algebra pdf 8
4.4 Free Vector Space according to unknown source 9
4.5 Free Vector Space according to "trans-la1" 9
5. Digression: equivalence classes and relations and related notation 10
6. The Wiki Tensor Product Page (including notion of Equivalence Classes) 10
7. My summary of what I think wiki is saying 13
8. Some other sources have their say: 14
8.1 Lang's Algebra book 14
9. Kronecker Products 16
10. Axler book has nada 17
11. Ganatra's class notes 18
12. Roman book preview 18
13. Roman Notes, Starting on Page 355 which is Chapter 14 Tensor Products. 20
13.1 The Universal Pair Concept 20
Universality Example 1: the triangle with inclusion map j on top 26
Universality Example 2: The quotient of a vector space divided by a subspace. 29
13.2 Direct Sums, External and Internal 30
13.3 Bilinear Maps 32
13.4 Tensor Products 34
13.5 Review of Ganatra and Explanation of why τ is linear 38
13.6 Conclusions 42
1. Tensor Doc use of tensor product in Appendix E
E.1 Direct Product Notation
This entire Section uses the general Picture A context where x-space need not be Cartesian,
(E.1.1)
The Standard Notation of Chapter 7 is used throughout. Useful forms can be found in Section 7.18.
The key tool required for the expression of tensor expansions is the notion of a direct product of n tensorial vectors defined in this simple way,
(ABC ...)abc... ≡ AaBbCc.....
(ABC ...)abc... ≡ AaBbCc..... etc . (E.1.2)
This is the Kronecker product, not a tensor product, and direct product is the wrong phrase!
The key tool required for the expression of tensor expansions is the notion of a direct product of n tensorial vectors defined in this simple way,
(ABC ...)abc... ≡ AaBbCc.....
(ABC ...)abc... ≡ AaBbCc..... etc . (E.1.2)
Again, this is a Kronecker product = outer product. It is not a tensor product or direct product.
The tensor ABC ... is nothing more than the outer product of vectors A,B,C as in (7.1.1) for contravariant vectors, but later extended to any mixture of vector types. One can define a direct or outer product of two rank-2 tensors in this way,
(MN)ab,AB ≡ MaANbB // rank(M) = rank(N) = n = 2; number of tensors = I = 2
(MN)ab,AB ≡ MaANbB etc (E.1.3)
Here at least I use the correct phrase "outer product, but I again sneak in the word "direct".
Notice how the indices are arranged on the left side of each equation. The same idea can be applied to form a direct product of tensors of any rank, for example
(MN)ab,AB,αβ = MaAαNbBβ etc // rank(M) = rank(N) = n = 3; number of tensors = I = 2
(E.1.4)
On the left side the number of groups of indices equals the tensor rank n of the tensors on the right, and the number of indices within each group matches the number I of tensors being direct-product-multiplied.
In what follows, only the (E.1.2) direct product of vectors shall be considered. One can define the dot product of two direct-product-space vectors in this obvious manner,
(ABC ...) (A'B'C' ...) ≡ (ABC ...)abc... (A'B'C' ...)abc
= AaBbCc..... A'aB'bC'c..... = AA' BB' CC' ... (E.1.5)
where of course the indices abc can be "tilted" in any way desired according to (7.11.3).
This seems like a reasonable dot product of two Kronecker products.
E.2 Tensor Expansions and Bases
Preamble: A seeming paradox and how it is resolved
This section does not mention tensor product objects.
Tensor Expansions
In what follows, we use the example of a rank-3 tensor and the reader can easily see how this applies to a rank-n tensor.
Let bi be an arbitrary complete set of basis vectors in x-space. As shown in the notes following (6.2.8) there exists a unique set of dual ("reciprocal") basis vectors bi (also in x-space) such that bi bj = δij, where we now use the Standard Notation bn equations of (7.18.6). Consider then the following expansion of rank-3 tensor A with contravariant components Aabc,
A = Σijk αijk (bibjbk) where (bibjbk ...)abc = (bi)a (bj)b (bk)c (E.1.2)
Aabc = Σijk αijk (bibjbk ...)abc = Σijk αijk (bi)a (bj)b (bk)c . (E.2.1)
OK, so what is happening here? I define a tensor in terms of the Kronecker product of basis vectors, not the tensor product. This does cause Aabc to equal the real number on the right.
As we did in Example 3 (2.9.5) for en, we declare the bn vectors to be "contravariant by definition" by writing the rule (b'n)i ≡ Rij(bn)j. So then the vectors bn are true rank-1 tensors under x' = F(x). According to Section 7.1, the outer product (bi)a (bj)b (bk)c is a true rank-3 contravariant tensor under F. Then (E.2.1) says that Aabc is a linear combination of these tensors, which we know is a tensor because the sum of two tensors of some type is a tensor of the same type.
The expansion coefficients αijk can be obtained by dotting both sides of (E.2.1) with (bi'bj'bk') and using (E.1.5),
(bi'bj'bk') (bibjbk) = bi' bi bj' bj bk' bk = δi'iδj'jδk'k . (E.2.2)
The result is then
A (bi'bj'bk') = {Σijkαijk (bibjbk)} (bi'bj'bk') = Σijk αijk δi'iδj'jδk'k = αi'j'k'
or unpriming indices,
αijk = A (bibjbk) .
But since both A and (bibjbk) are elements of the triple direct-product space spanned by the vectors
(bibjbk), we use the dot product of (E.1.5) to claim that
A (bibjbk) = Aabc (bibjbk)abc = Aabc (bi)a (bj)b (bk)c .
The coefficients αijk may then be written in all these ways :
αijk = A (bibjbk) = Aabc (bibjbk)abc = Aabc (bi)a (bj)b (bk)c (E.2.3)
This "space spanned by basis vectors (bibjbk) now sounds more like a tensor product. Ouch!
where Aabc are the contravariant components of tensor A in x-space, and (bi)a are the covariant components of vector bi in x-space. As noted in Example 3 of the previous subsection, the object αijk transforms as a scalar, since it is the dot product of two direct-product-space vectors. This fact is especially obvious from the last expression in (E.2.3) where all tensor indices are contracted so the result must be a scalar according to the neutralization rule (7.12.1). So αijk is a set of 33 =27 scalars.
OK, let's stop here. It seems clear to me that I am only talking about the Kronecker = outer product in this appendix.
Then we have
E.4 Dyadic Products
When two vectors A and B are combined in polyadic notation, the result is called a dyadic product (AB) [ also known as a dyad or just a dyadic ]
(AB)ij ≡ AiBj // = (AB)ij = component of the matrix AB = AB . (E.4.1)
So again I am talking outer product, not tensor product. Except for my use of the word "direct", everything seems to be self-consistent here, and a tensor really does seem to be an expansion on the outer product of the basis vectors. In particular, the vector is this
V = Σiviei Vk = Σivi(ei)k
so even this simple "tensor" can be written in terms of those real basis vector components.
__________________________________________________________________________________
2. Connection between tensor doc and my current tensor thread.
So how does the above self-consistent tensor doc stuff relate to my new tensor product discussion where I really seem to use the tensor product and not the Kronecker product? That is the $64 question! I write things like this:
{ei} = basis of V v = Σi=1n vi ei = general vector in V
T ≡ ΣijTij eiej . = general object in V2 T ϵ V x V = V2
These two objects exist in V1 and V2, for example. {ei} is an arbitrary basis like {bi} of tensor doc.
I never talk about "components" of something like eiej, though I do mention (ei)k in a few places. In tensor doc, I would say
A = Σijk αijk (bibjbk) where (bibjbk ...)abc = (bi)a (bj)b (bk)c (E.1.2)
Aabc = Σijk αijk (bibjbk ...)abc = Σijk αijk (bi)a (bj)b (bk)c . (E.2.1)
and I would claim that A was a tensor expanded on the {bi} basis having strange components as shown.
This is indeed very confusing right now!!
Can you have something that is both a tensor product and, within that, an outer product as a sort of subclass of tensor product? Look again at wiki on tensor product. Wiki uses (v,w) notation but writes the space as VW. They have a very good comment on quotient space which I defer for the moment.
4. A Free Vector Space
Since this concept is used in the Wiki Tensor Product discussion, I will have a section right here to examine this sub-topic before moving into the Wiki stuff.
4.1 Free Vector Space according to the wiki tensor product page
Here is the rather confusing wiki paragraph on free vector spaces. Comments will come after:
In this following clip, wiki states its definition of "free vector space" F(S) as f: S→ K(a field). Notice that the F and the K are not the same. The F means "Free".
The nature of the set S is not clear to me. The clip refers to a "finite subset of S". I take that as meaning a subset having a finite number of elements, not a subset that is bounded. The clip says that S, in addition to being a set, is a vector space, which means it has addition and scalar multiplication amongst its elements s. Could we select S = R, the reals? That set is a vector space, but how do you pick a finite subset of S? I guess I am wondering if S can be a "continuous set"? I think the answer is no. The idea of δs(s') = δs,s'can be extended to continuous sets as the Dirac, but I am pretty sure S is a discrete set.
The clip says any element of F(S) can be expressed as a linear combination of those δ's. That seems to lead to the following rather uninteresting set of functions
f(s') = Σsfsδs(s') // "linear combination of the δ's
= Σsfsδs.s' = fs'
So, what is the meaning of the most general function of F(S) having the form f(s') = fs' ? This seems to say nothing at all, like 2 = 2.
If S cannot be continuous, that we cannot say for example that S = R3 with basis elements ei. Could we instead say that S = {ei} for i = 1,2,3. But such an S is not a vector space, it is the basis of a vector space.
Conclusion: Although wiki provides some info on the notion of a free vector space, the full picture of such an animal is extremely unclear. What sets S are really allowed? So we shall now look at some alternate web offerings on the subject of free vector space.
4.2 Free Vector Space according to Drexel
https://drexel28.wordpress.com/2011/04/19/free-vector-spaces/
Instead of a finite subset of some possibly infinite set S, Drexel starts us off with a finite set X.
Instead of talking about F(X) = a set of functions f: X→ reals, Drexel talks about F[X] = a set of sums of the elements of set X weighted by reals. This is I guess a special case of a function f. So Drexel would have his most general F(x) element looking like this,
f = Σx fx x = element of F[X]
as an element of his F(X). In a later comment he suggests that you could label your x's in X as xi and then the above becomes
f = Σi fi xi = element of F[Ix]
So this is still a sum over X, but we have a certain way to label the x's (I is an "indexing set"), so here you could consider S = I or S = X, your choice of viewpoint.
He does NOT say that the x or xi are real or in field F, though he might intend that.
He does NOT say that X is a vector space (whereas wiki says S is a vector space). But perhaps I could just make X be a vector space by having it be closed under addition etc.
The two "rules" he give don't really seem like rules. The first says,
f + g = (Σi fi xi) + (Σi gi xi) = (f+g) = Σi (fi+gi)xi
Since the sums are finite, the last item is just rearranging the summation terms, so this is not even a rule.
The other "rule" says
α*f = α* (Σi fi xi) = Σi (αfi) xi
Again, since fi and α are reals, this does not seem like much of a rule!
Why is F[X] defined in this way a vector space? Well, the space of sums of the above form is indeed closed. There is a zero sum for use as the "identity" element. Inverses exist and so on.
So I will state his definition in my own words:
" A free vector space over a finite set X with elements xi is the space of sums of the xi with real coefficients. " OR " F[X] is the free vector space of X over R ".
So Drexel's definition allows me to think of the xi as my basis vectors of R3 which are ei. These ei are not reals, they are tuplets of reals, but Drexel does not require that the xi be reals. I'll be you are supposed to think of the xi as "inert carriers" like the x in a polynomial. Drexel goes on to say:
This would then give dim = 3 if I choose X = e1, e2, e3. Note that x is called a "formal object", like my idea of inert carrier.
4.3 Free Vector Space according to Algebra pdf
This is closer to wiki in that we have f : S → R with its same idea of finite support. Whereas Drexel had only sums for functions, here we have general functions, again like wiki. But now we have rules on these functions which are the same rules that Drexel gave for his finite sums. The first rule just says this free vector space is closed, the second shows that (αf) is a certain function, so you can scale functions. The elements of set S can be any objects you want, but the functions f(s) have to map to reals. You could I suppose have S = {ei}, but not clear what typical functions f would look like. Author claims that S is a basis for F, the free vector space. This offering fits roughly with the previous two in its outline, but the specifics are vague.
4.4 Free Vector Space according to unknown source
The set X is not specified to be finite.
The object F is not defined, so unclear what it means to have a direct sum of undefined objects!
The basis functions are similar to the δs ones of wike, where now s,s' replaced by x,y.
Maybe the direct sum means the space of sums like those of Drexel with coefficients from F.
Σx fx x
Not very helpful.
4.5 Free Vector Space according to "trans-la1"
So this is a mapping from integers to reals (let's say) fi = f(i) a real number. It could be a sequence of reals. The values of i are in an index set I, fine. Why is this a vector space? What are the vectors of this vector space? I guess the vectors are the functions f. So we have to show that
α(f + f') = αf + αf' etc
but since the f's map to F, this all follows naturally. Now author continues,
Here the set S is finite set of integers from the integer set I.
The functions f in effect map from I to R. No "inert carrier" is specified so on xi here.
Otherwise, this is little like the Drexel 4.2 approach.
So we have five different definitions of a "free vector space".
Roman mentions "free vector space" in his Preface, but that phrase never appears in Chapters 14 and 15. So he is not very helpful.
5. Digression: equivalence classes and relations and related notation
Stop. What is the equivalence class of a set S? OK, wiki says this is a concept of set theory where you just partition your set in to subsets in each of which elements are equivalent by some rule a ~ b . The rule is not specified.
Sometimes this is written, where a ~ x is written a R x and R is the equivalence relation
[a]R = {x ϵ X | a R x }
So far so good. Suppose you successfully partition your set using R so you get
S = { [a], [b], [c] } which has 3 equivalence classes for some relation R
The equivalence classes themselves form a set which is known as X/R " X modulo R" "quotient set". Another notation for this set is then X/~ which I have seen!
6. The Wiki Tensor Product Page (including notion of Equivalence Classes)
" Notation
Elements of V ⊗ W are often referred to as tensors, although this term refers to many other related concepts as well.[1] If v belongs to V and w belongs to W, then the equivalence class of (v, w) is denoted by v ⊗ w, which is called the tensor product of v with w. In physics and engineering, this use of the "⊗" symbol refers specifically to the outer product operation; the result of the outer product v ⊗ w is one of the standard ways of representing the equivalence class v ⊗ w.[2] An element of V ⊗ W that can be written in the form v ⊗ w is called a pure or simple tensor. In general, an element of the tensor product space is not a pure tensor, but rather a finite linear combination of pure tensors. For example, if v1 and v2 are linearly independent, and w1 and w2 are also linearly independent, then v1 ⊗ w1 + v2 ⊗ w2 cannot be written as a pure tensor. The number of simple tensors required to express an element of a tensor product is called the tensor rank (not to be confused with tensor order, which is the number of spaces one has taken the product of, in this case 2; in notation, the number of indices), and for linear operators or matrices, thought of as (1, 1) tensors (elements of the space V ⊗ V∗), it agrees with matrix rank."
So this does suggest that the outer product might be a subclass of the tensor product, but I need then to know more about "equivalence class". I do like their claim that physics really does use the outer product, which is what tensor doc uses. Now look back at the opening wiki paragraph,
"In mathematics, the tensor product, denoted by ⊗, may be applied in different contexts to vectors, matrices, tensors, vector spaces, algebras, topological vector spaces, and modules, among many other structures or objects. In each case the significance of the symbol is the same: the freest bilinear operation. In some contexts, this product is also referred to as outer product. The general concept of a "tensor product" is captured by monoidal categories; that is, the class of all things that have a tensor product is a monoidal category."
Notice the key word "bilinear" and "freest bilinear". I have not looked at all into "monoidal".
So let's back up and try to parse wiki's page on tensor products.
I like using K for the field, all is well so far.
Now, here is where the "free vector space" buzzword appears, and this is why I tried to find someone who have a clean definition of such a space (see above).
Note that the "free vector space" F(VxW) is the set of functions f such that f : VxW → K (field) and which have finite support over VxW. Such a function would be written r = f((v,w)) where think of s = (v,w) = element of S.
Note 1: I understand V x W as the Cartesian product with elements (v,w). Take the (v,w) elements to be the set S. This would seem to define a continuous set S. They talk about F(V x W) so are taking S = V x W. Well, the elements of set S are things of the form (v,w) so there are a lot of such points in S if V and or W are continuous sets, yes.
Note 2: If we apply the Drexel idea, then F(VxW) is a set of functions which are sums of this form
f = Σi fi (vi,wi) = Σij fij(vi, wj)
where vi is some arbitrary element of V, and so on. Here for example is a possible sum, for some v and w
f = c(v,w) - (cv,w) .
Again, sums are special cases of functions, so the wiki thing of 4.1 is OK too. This last function is going to get forced to zero soon because it is in an equivalence class with 0.
Now back to tensor product. I know what the equivalence classes of S means, so then if S = V x W I have to ask: what are the equivalence classes for the set elements which have the form (v,w) ? Whatever these classes are, they are used as the vectors of U = VW. It seems that the scalar and distributive rules are imposed on the (v,w) as an extra feature.
Now here is another piece to chew on:
They are talking about the meaning of + and * for elements of U ≡ VW where * is scalar mult of an element of U. I guess 1 is a "representative of an equivalence class" of S = V x W, and then maybe the class itself we can call u1. The next lines say this:
For two representative elements of two equivalence classes (which classes are elements of U), you can do addition in this way: 1 +F 2 where +F is the appropriate addition operator for U. It is what makes U be a vector space. Then to do scalar multiplication you have c F 1 where F means scalar multiplication of an element in U. Any representatives will do.
So one idea is that you don't add and scalar-multiply the classes directly, you add and scalar multiply elements of those classes. I am fine so far. (familiar from Galois doc I think)
Note Added: This equivalence class stuff is just a fancy way to say that the + and operations you have in the space VW respect the "bilinear rules". It is I think better stated in the following clip:
A whole new batch of stuff here. What is this little V symbol on the right? It is the OR operator in logic notation, fine. So at least one of the four equations has to be true.
What then is set N? First, what is F(VxW)? This is the set of functions f: VxW → K so these functions are K-valued. A typical function might be a(v,w) + b(v'w') where I guess + is addition in field K, and where a and b are elements of field K. So F(VxW) is the set of K-valued functions of (v,w) elements.
Note: Just maybe f: VxW → K somehow means that the (v,w) are my "inert carrier" objects and the → K just means that the coefficients are in K. It is then more like f: VxW → T where T is the space of linear combinations of the (v,w) inert (formal) objects with real coefficients. In this context, we are going to knock out from T certain elements.
Now four such functions appear on the right of the above 4 equations. Obviously if the scalar and distributive "rules hold", n = 0 for each of these four functions.
So I guess N is these specific four functions! Each of these functions is an element of F(VxW), I agree. I can see that N is a subset of F(VxW). There are really more than 4 functions because you get a function for each choice of v1, v2 etc. No doubt if you add two of these functions, you get a function in the same set, etc etc, so I suppose the set N could be a subspace of F(VxW).
Now using our rules, all these functions are equal to 0, even though they are all different functions. That is to say, every function in N equals 0 which is an element of field K I guess. Are such functions in the same equivalence class? Depends on how you define your relation ~.
I guess you could partition F(VxW) into classes where in each class the function takes some value in the set K. Then since 0 is an element of K, all the N functions above are in the equivalence class of functions which equal 0. I guess there might be more functions that those shown above which = 0. Maybe we define N to be an equivalence class and we label that class by the name [0]? Again, N is an equivalence class whose elements are n and such a class is written [n]R. Maybe there are only two classes, N and all other functions in F(VxW). The notation ϵ F(VxW) seems strange. It is VxW = set S which we partition into equivalence classes, so VxW = [0] + other classes. Well, I suppose if you were to partition S into classes so S = [α] + [β], you could talk about F = F([α]) + F([β]) and maybe these are then induced classes for F(S)?
OK, now comes the big line: VW is the quotient space F(VxW)/N where N is the subspace of F defined above. I guess you can take a quotient in this manner.
7. My summary of what I think wiki is saying
1. The free vector space F(VxW) is the set of all finite linear combinations of (v,w) formal objects with real coefficients. These linear combinations are the Drexel sums above. If you were to single-index the (v,w) elements, you might call them (vi,wi) and then a most general element of the free vector space F(VxW) would be this
f = Σi fi (vi,wi)
Here fi are real coefficients, but the (vi,wi) is "formal" and not of any particular type. It need not be, for example, that (vi,wi) is a real number. It has no specified type IMHO.
2. Certain of these functions, such as f = 2(v,w) - (2v,w), are forced to be zero. This is done by saying for example that 2(v,w) - (2v,w) ~ 0, meaning this function "is equivalent to 0". Certain functions within F(VxW) are by fiat made to be equivalent to 0 in this manner.
3. Consider the following function which is a function in the free vector space F(VxW)
f = 3(v,w) + [α(v,w) - (αv,w)]
As one varies α, one has a lot of functions represented by this equation. These functions are elements of a row of the "chart" for the quotient F(VxW)/N . We can take the row representative to be 3(v,w). The row of the chart is an "equivalence class". All its many elements are represented by 3(v,w). So in the object that is F(VxW)/N, the only elements are the chart rows! This is a vector space whose elements are the rows of the chart. And 3(w,v) is a chart row, so is in F(VxW)/N . Then if we say
VW = F(VxW)/N
we are saying that the elements of the vector space VW are those chart rows. In this VW space, the quantity α(v,w) - (αv,w) is equal to 0.
4. In less formal language, VW is the set of all functions in F(VxW) where we force certain of these functions to be 0. Just as an element of VxW is written (v,w), an element of VW is written vw. Then similar to the above
f = 3vw + [α(vw) -(αv)w] = 3vw
where we have set the function α(vw) -(αv)w = 0 by our definition of the VW definition.
5. You could say that VW starts out as VxW, but then we overlay certain "rules" on linear combinations of the (v,w) and when such rules are observed, we refer to (v,w) as vw. There rules are I think exactly the "bilinear rules".
8. Some other sources have their say:
8.1 Lang's Algebra book
Regarding tensor product vs other things, Lang's "Algebra" supposedly has material. It is this book, there is one chapter of great interest
The book is 934 pages or so. I found a very slow Chinese server, trying to load it in. Have it!
But alas, it is densepack jargon of the worst kind for me! I would have to spend 6 months to be able to ready the chapter, and I don't really think it has what I want. But at least I have a nice book.
Note added: Well I reread it, and except for the fact that he uses the "module" word a lot (remember that a module is a vector space over a ring instead of over a field), he does exactly what I have just outlined above, concluding that VW = F(VxW)/N. But Lang shows that
E1E2......En = F(E1xE2x......xEn)/N.
where he does more than two items. But his approach is more like the Roman one below where we use those triangle diagrams and the universal buzzward. Lang says:
"We use the words linear and homomorphism interchangeably".
and here is his picture (notice the appearance of "free module" rather than "free vector space") And he suggests that the space N is a subspace of M, though that word did not appear in the above stuff.
I look more at the wiki refs and see Lang's Algebra sitting there, and Birkhoff Mac Lane. Another one is a book of Bourbaki which I just loaded in. Have it, all about rings and modules and fancy things relative to the tensor product. Lots of Hom operators, way too advanced for me.
Hom(E,F) = set of linear mappings E to F. // see Lang's comment above
I am looking for more of a primer on my topic, not the grand master work.
9. Kronecker Products
Here from https://en.wikipedia.org/wiki/Kronecker_product :
From this definition, the Kronecker product is the generalization of the outer product of two vectors to the outer product of two matrices, as per my tensor doc statement that (ST)ab,AB ≡ SaATbB
But here is more from the same wiki page:
Well this is that QM notation of the cross product of two operators in different spaces
The claim then is that
(ST)(vw) = (Sv)(Tw)
(ST) : VW → (SV)(TW) = XY
Suppose I take v = ei and w = e'i. Then the first line says,
(ST)(eie'j) = (Sei)(Te'j) .
Now I know that Sei is a vector, so lets add some subscripts to the action
[(ST)(eie'j)]ab = [(Sei)(Te'j)]ab = (Sei)a(Te'j)b
Now consider
(Sei)a = Sar(ei)r = Sar δi,r = Sai similarly (Te'j)b = Tbj
Then we end up with
[(ST)(eie'j)]ab = SaiTjb
Then I would have to define
(eae'b )T(ST)(eie'j) ≡ (ST)ai,bj = SaiTbj
and that then produces the desired result. The left side could also be written
(eae'b )T(ST)(eie'j) = (eaT S ej)(e'bT T ej') = Saj Tbj
which is another way to get the result.
A different take on this goes back to
(ST)(vw) = (Sv)(Tw)
where this gives the quantum mechanics Hilbert space notion that S acts on the first item and T acts on the second item. This is a tensor product operator acting on a tensor product vector, so is really a new concept from the regular tensor product idea.
Possible link
https://www.dpmms.cam.ac.uk/~wtg10/tensors3.html
it is very wordy and has poor graphics but it does talk in primer language.
I have looked pretty hard on the web under "tensor product" "outer product" and I just don't see anyone talking about the latter being a special case of the former.
Most relevant sources do talk about a "free vector space" defined on sets.
10. Axler book has nada
Sheldon Axler's Linear Algebra Done Right.
Try to get this item: got it, but it has nothing, it is just a basics LA book.
11. Ganatra's class notes
I just invested heavily in this guy Ganatra,
http://math.stanford.edu/~ganatra/math113/
In homework 7 you are supposed to show a certain triangle fact, and he even has a solution set, but I don't think the triangle fact is proved. You have to show that something is linear if two other things are bilinear. Basically, no proof is given in the homework. But the buzz word "universal bilinear map" is mentioned and maybe that would lead to another source on this.
https://books.google.com/books?id=bSyQr-wUys8C&pg=PA361&lpg=PA361&dq=%22universal+bilinear+map%22&source=bl&ots=Laf7MeHJWr&sig=LUDaYqVMPeG6jGW_3T2goLvUtuU&hl=en&sa=X&ved=0CDgQ6AEwA2oVChMI1MTigM38xwIV2A-SCh39CgCo#v=onepage&q=%22universal%20bilinear%20map%22&f=false
Here is some similar material: [ this is where I found the Roman book probably in google books ]
12. Roman book preview
Now clear why this is a "guide for the definition of tensor product U V". I think perhaps the idea is that there is really only one bilinear map off UxV. Here they show one going to T and another to W, but since T and W are linearly related, perhaps the bilinear maps are in some sense the same. Let's try do decode some of this guy's stuff, I doubt I will get very far.
Decoder activating: hom shown is the set of all bilinear maps from UxV to ANY vector space W, so I guess S is a set of candidate vector spaces for W. But I don't know what "measuring family" means. Maybe he is defining: "measuring family" ≡ all linear transformations in the universe??
Now how obscure can he make it:
All these glossy words just say that the τ link in the triangle is linear. I guess T and t here are two maps of the same type. The "pair" here is the space T and the mapping t. The pair is what is universal IF the rest of the statement is true. This is just a definition of universal, there is no proof that anything is linear.
So T = UV is a tensor product if {t,T} is a universal pair. Every universal pair defines a tensor product. Again: Any universal pair {T; t:UxV→T } defines a tensor product UV ≡ T. So I guess if you want to test some specific UV for being a tensor product, you have to identify the "universal pair" thing.
The text then skips two pages. The book here is
Advanced Linear Algebra, by Steven Roman (in google books). Can I find that book? I got the TOC at least. OK, found a slow copy from China at 5.23 MB, then I can continue with the above presentation. It would be good if the author would give AN EXAMPLE!!! Can you always find a universal pair? I guess you have to say something about the space T.
OK, I have the Roman book and it is very good! I need to study it a bit in my next session. I think it makes the most general definition of a tensor product, and I think I will be able to show that the outer product is a particular case of a tensor product.
13. Roman Notes, Starting on Page 355 which is Chapter 14 Tensor Products.
13.1 The Universal Pair Concept
"Universal Pair" is a concept in "Category Theory". Roman will use "simple terminology".
Notice that nothing is bilinear or linear yet, just a generic picture.
So the ability to "distinguish" elements in A is important, and the above triangle for me says that
f(a) = f(b) g(a) = g(b), so the contra shown above. So f and g have "something in common". Notice that no statement is made about the nature of mappings f,g,τ.
Injective means 1-to-1, so then we can say f = τ-1 o g and then we get the other directions above, and then f and g exactly distinguish things in A in the same manner.
Set S seems to be the set of all possible spaces X (no claim of vector space). Then S is one of those spaces so we have g: A→S. Note that "g: A→S" is an element of F. Imagine there is another element of F which is "f: A→S" (F is a very large set of mappings from set A). In some sense, f and g are both functions associated with this space F. We just said above that for any S, the information about A contained in g is also contained in f. By letting g and f wander in F, I guess all the functions in F have the same information about A as does a particular function f, assuming the above triangle idea. I would perhaps say that f was "representative" of F, and as such, contains the same info about A as any other representative. Their word is to say that f is "universal" in its information content, all other g have the same information. The real phrase is that "f is universal among the functions in F.
The pair here is (S, f), a particular range space for function f, and the function f. I guess we can apply this idea to any set A, so we don't need to indicate A in this notation. I would like to see some examples. Not obvious to me why a basis for a vector space is an example.
So F is exactly as before, but the new item is H. In the first thing, X is some space in the set S. In the second thing, this space X is the domain, and the range is some other space in S . Also part of H is a new function τ. Basically, H is the space of all mappings between any two spaces in S . But we now add requirement on the set of functions H:
What does that mean exactly? Well, that applies only if X = Y, and then iX : X → X could be a well defined identity function, so that is OK.
I think this means that if τ1: X→Y and τ2:Y→Z, then τ1 o τ2: X → Z "makes sense", where X,Y and Z are all in S. I guess this would be "associative" because for more functions o is associative?? Unclear.
OK, this is just a straight assumption, fine by me. Here then is a picture:
This fancy picture is one of many you could make. Three arbitrary spaces Si as shown, and A. The three functions fi are all obviously in the set F . And the three τi functions are in H. Notice how assumption 3 really says that τ1o f2 = f1 and also τ3o f2 = f3 and also τ2o f1 = f3. Meanwhile. assumption 2 allows the fact that τ2o τ1 = τ3 . I write these equations using "commuting paths" in the picture.
But what is it that they are measuring?
Definition of a Universal Pair:
In this chunk, Roman is zeroing in that universal buzzword. Here is what I think the defining clause says starting with "if for every": for every (g,X) and (f,S) in F, there exists a τ in H which is unique AND which allows the above picture to commute, which means you can write g = f o τ ("g can be factored through f"). If this situation is true, THEN
The pair (S,f) has the universal property for set F as measured by H
The pair (S,f) is a universal pair for (F,H )
In the picture, the unique τ function is called the "mediating morphism for g".
OK, look at the above picture where f and g are in F . If the triangle picture can be drawn, then (S,f) is almost a universal pair. The only extra thing you need is that τ is unique.
Recall from earlier category notes and a morphism is an arrow in a category diagram like that above. It is also a map which "preserves some structure". Here, perhaps that is the information-bearing capability of the functions f and g. Perhaps it is just the commuting of the diagram is the structure being preserved.
Unclear. The universal pair is the thing (S,f), so how can this pair be unique? (X,g) is another pair that is a different pair. Yes, they are linked by a unique function τ. We continue on:
Here we have two pairs (S,f) and (T,g). Each is a universal pair. I don't know what μS=T means. Perhaps for all elements s in S, you get μ(s) = t where t ϵ T. Also unclear: if τ1 and τ2 are the two unique functions as indicated, how can two functions be an isomorphism? Maybe if you run over all f and g in F, the set of τ1 mediating morphisms f to g is isomorphic to the set τ2 for g to f. Hopefully the following proof of the theorem will help to clarify what it is trying to say. Proof starts off,
On his second line, he is saying f = σ o (τ o f) = (σ o τ) o f which quietly makes use of the assumed associativity property of the composition operator o. So far all is fine. We continue,
OK, in that third picture the unique dotted line function is (σ o τ) since we know f = (σ o τ) o f. On the other hand, we know that f = iS f where iS is the identity operator. Since things are unique, we know then that σ o τ = iS. A similar argument would show τ o σ = iS Thus, σ and τ are inverses of each other (left and right inverses). So the function τ is in fact 1-to-1. Not exactly clear why τ is "onto" but I guess it is defined for every element in S, and therefore it is a bijection as stated.
I still don't see the significance of the universal pair concept, nor do I see what is "unique" about it. I guess in the left picture, Roman is studying the universal pair (S,f), while in the middle (T,g). Since it turns out that τ and σ are inverses, maybe THAT is why one says that (S,f) and (T,g) are essentially the same, and so in that sense there is only one universal pair out of A, and so that pair is then unique.
Universality Example 1: the triangle with inclusion map j on top
Vect(F) = all vector spaces over field F. Note terminology I have often wondered about
family = class = collection that is too large to be considered a set
There are certainly many infinite vector spaces on the reals, for example.
Let's review the soup ingredients:
B = a set of at least one element (it will be the set of basis elements of a V)
S = all those vector spaces Vect(F) . Before, this was just a collection of sets.
F = the collection of "set functions" f such that f: B → F
so this function f maps elements of the set B into other functions in F, strange.
H = set of linear transformations I suppose from any space to any other space.
Now the idea is to take as the set B a set of basis elements of vector space VB.
But what is an "inclusion map"?
The inclusion map takes things in a subset and maps them into the larger space. So we can then for example map those basis elements from the smaller set B into the full vector space VB. That is what the mapping j does. The pair under study here is then (VB, j) or in more detail (VB, j:B→VB).
Now somehow Roman wants to say that this is an example of a "universal pair". We have to show that there is that triangle picture with a unique τ.
Repeat text above
What is object v here? It must be a basis element in B. We can map j(v) = v into VB, then apply τ to get τ(v) which he writes as τv (a notation unclear to me). But start with the same v back in B and then g(v) = gv. Thus τv = gv. Write again as τ(v) = g(v) for all v in B. This seems to say that τ = g, and then certainly τ is unique, given g.
Now why is τ linear? You would have to show that τ(v + v') = τ(v) + τ(v'). That is the same then as g(v+v') = g(v) + g(v'), but Roman has not specified that g is a linear mapping. I keep losing the thread in this kind of discussion when it comes to τ being linear! This is really the first appearance of the word linear, except he imposed that H = linear transformations.
The only way out I can see is this: Roman must have neglected to say that g ϵ F where F is a space of linear functions. Then when you end up with τ = g, then τ is also linear. He continues,
I continue to think that g(v) must be linear. The above paragraph does not help me much. I hope this confusion won't wreck the rest of what Roman has to say.
[ Note added 9.18.15 7 PM: Go back to the above picture
where we have
j(v) = v in VB inclusion map
τ(v) = x
x = g(v)
Then we can define τ by
τ(v) = g(ν)
Now suppose we simply declare that τ is linear, as I learned much later in these notes. Then we have
τ(v + v') = τ(v)+τ(v')
But this says
g(v+v') = g(v) + g(v')
Therefore, if you want τ to be linear, you must also have g be linear. ]
Universality Example 2: The quotient of a vector space divided by a subspace.
Next example, here is the setup:
OK, now he says that F is linear maps. Also, K is a subspace of V and also this same K exists within the nullspace of any of those linear maps in F. So this is a much more restricted collection F . For example, if f is in F and if k is in K, then f(k) = 0.
Now I need to go access Theorem 3.4 from earlier in the book: I leave p 366 and go to p 104
Note that L(V,W) is defined on page 73 as the obvious
So in Theorem 3.4 and its triangle picture, τ is a linear transformation τ:V→W. And in this theorem S is a subspace of V and all of S lies in the kernel of τ, so τ(s) = 0 for any s ϵ S. But now I need to know the meaning of the notation V/S. Well, in Galois doc I first discuss G/H being the factor group when H is a subgroup of G. Later I discuss R/I as the residue class ring for an ideal I in the ring R. But I never talk about the notion of A/B where B is a subspace of vector space A. That topic is addressed by Roman on page 100 and it is very similar to the group discussion. You write [v] = v + S, meaning that all elements v+S are in the same row of the chart (v = representative) because here f(v+s) = f(v) since s is in the kernel of all functions f of interest. So the set of rows [v] form the quotient group which then is V/S.
So not dwelling too much on those V/S details, Theorem 3.4 above is saying that if πS projects elements of vector space V onto those "chart rows", that is, into V/S, then for any linear τ:V→W, you can find a linear τ' such that τ' : V/S → W and you will then have τ' o πs = τ.
This certainly seems a strange little world. Why would you want to carry out mapping τ by "factoring through πs" and thus passing through the quotient space V/S of V relative to S on the way from V to W ?
Comment: since f(v+s) = f(v) for all s in the subgroup S, there is sort of only one f(v) here and you might as well talk about f([v]) and not bother with a list of all the s's. So you save effort by talking only about f([v]). You "mod out" all the equivalent arguments. So that is a good reason to work on V/S.
But let's just accept this rather obscure Theorem 3.4 and get back now to where we were.
The pair here is (V/K, π) and we can then way in our local notation,
Since τ is a unique and linear transformation in H, you can conclude that (V/K,π) is a universal pair, and that is what this is an example of. I guess π is called a canonical projection map in that earlier section.
[ Why is (V/K,π) a universal pair: (1) because the above triangle exists; (2) τ is linear. ]
I now skip example 14.3 regarding direct sums being an example of universal pairs. I hope it won't be needed for what I am trying to track down by reading in Roman.
13.2 Direct Sums, External and Internal
What is this saying? When he says vi ϵ Vi, I guess vi is a vector which has a number of components equal to the dimension of Vi. So I think the element of the direct sum space is just a vertical stack of vectors of different sizes, and this is what I have always thought of as the direct sum. A matrix in block diagonal form would then map such an stacked element of V into V with no mixing. I have always used the symbol for this direct sum idea, but he uses a boxed plus sign.
But in the above, the Vi are ANY vector spaces. In Roman's "internal" direct sum, it is required that each of the Vi be a subspace of V. So then he writes V = V1V2 ... using the symbol. So this then would be a special case of the box-plus symbol direct sum. Now back to where we were:
Consider the meaning of L(V x V, W) whose representative reads w= f((v1,v2)). The domain element is simply the pair (v1,v2), an element of V x V. Linearity (implied by L) would mean
f( (v1,v2) + (v'1,v'2) ) = f( (v1,v2) ) + f ((v'1,v'2))
Note that there is no sense of "bilinearity" here. That is, we do not have,
f(v1+ v2,v3 ) = f(v1,v3) + f(v2,v3). for example (part of bilinearity)
We cannot extract either argument of a pair in the first line and do something with it.
Now go back to the first concept with f( (v1,v2) ) . Think of (v1,v2) as a stacked vector, then you would say (v1,v2) = element of both V x V and V ⊞ V. You have a direct sum structure (external type). Probably V is not a subspace of V x V, so you do not have the type of direct sum V V here. Remember that V ⊞ V is the rawest most generic direct sum.
I think the above quoted says: start with the Cartesian thing V x V and you have as yet no algebra. But then add the set of linear maps L(V x V, W) defined on V x V, then you have some linear functions going on, f: V x V → W. This is the first of two possible structures you could have. The second is that you define a bilinear map w = f(v1,v2) and then your structure is this set of bilinear functions on V x V. Recall that this set of bilinear maps is called hom(V,V;W) whereas the first is sort of V ⊞V.
13.3 Bilinear Maps
New Roman Topic
He then defines bilinear maps and I am fine with that. Here then are some definitions
Notice these definitions and notations:
set of all bilinear functions f such that f:UxV→W is called homF(U,V; W).
I think the F refers to the field used for scalars in the linearity stuff, called "the base field".
This set is in fact itself a vector space. Probably there is some reason for the name "hom" relating to homomorphism, but no one wants to explain it on line. // Well above I quoted something that said with vector spaces, homomorphism and linear mean the same thing.
If f:UxV→F, that base field, then f is called a bilinear form.
If f:UxV→R, I have been referring to this as a bilinear functional.
So elements of U x V are just (u,v) and a priori they are not part of any "algebra". I agree.
Now here is an example of the bilinear case.
I think of A as a set for which multiplication is defined. Then f: A x A → A says μ(a,b) = ab and since this is a product, it is obviously bilinear. Reminds me of the outer product idea yet to come.
Again: this Example 14.4 is an example of f : V x V → W with the bilinear structure. It happens that W = V so we have μ : V x V → V with w = μ(a,b) = ab. Manifestly bilinear!
I skip Example 14.5, and go to
Here f and g are linear functionals, and this causes φ(u,v) defined above to be bilinear!! The Dually part above is certainly just a rehash of the first case with V* and W*. The functionals appear in the same way in both examples. I am happy with this example, both parts.
Notice this statement that your basic skeletal tensor product is bilinear and nothing else!
I have now done all the background work and am ready to go over again the upcoming section on tensor products. On my previous pass, the background was not there yet.
(Resuming here at 3 PM on 9.18.15)
13.4 Tensor Products
I am just quoting text as it appears, now on Roman p 362. Nothing more is said at this point about the above picture. For example, no proof is given that if f and t are bilinear, then τ is linear. But I will take a quick shot right here:
Plan A:
t1 = t(u.v) is bilinear
w = f(u.v) is bilinear
what then do we know about τ which appears in w = τ(t1) ?
To show that τ was linear, I would have to show that
τ(t1+ t2) = τ(t1) + τ(t2)
or
τ( t(u1,v1) + t(u2,v2)) = τ( t(u1,v1)) + τ( t(u2,v2))
or
τ o ( t(u1,v1) + t(u2,v2)) = τ o t(u1,v1) + τ o t(u2,v2)
I know that
RHS = f(u1,v1) + f(u2,v2) = w1 + w2
So I would have to show that
τ( t(u1,v1) + t(u2,v2)) = w1 + w2
But if I know nothing at all about τ, then I cannot write τ( t(u1,v1) + t(u2,v2)) in any way!
Plan B:
Try writing τ o t = f as τ = f o t-1 where t-1: T → U x V. But it seems to be very unlikely that t is invertible, so this really goes nowhere.
Comment: Having tried this about three times now, I am getting the idea that maybe linear τ is by fiat and not something you prove. It is part of that universal buzzword definition. If you have τ linear, then maybe you can show also unique and then you have universal pair (T,t). Let's just hack onwards in hopes of getting some answers in a few days.
[ Note added 9.18.15 7 PM.: I finally show below that you do in fact just declare τ to be linear as part of its definition. When you do this, you check to make sure everything in the triangle is consistent! You can never prove that τ is linear by Plans such as the above! ]
Roman continues (no gap)
It seems that τ linear really is being imposed or assumed here by saying H is linear [correct]. But then all he really says here is that the pair (T,t) is universal for functions in set F IF you can find a unique and linear τ. So Roman is just repeating earlier material in this bilinear space context. Again, the above entire quote is just a definition of what you would have to show to show that pair (T,t) was universal for F . Nothing at all has been proved or even claimed. We continue on some more:
OK, this defines a tensor product in terms of (T,t) being a universal pair for the space F of bilinear functions. So if you can somehow show that a particular (T,t) pair is universal, THEN you have a tensor product. So how are we to show such a thing if we are handed a candidate tensor product?
OK, we continue onto page 363
OK, I am all for "construction". Continue on:
OK, above he is writing
u v = Σij ui vj( ei fj) u = Σiuiei v = Σj vjfj
which is something I write all the time. If you defined u' = Σiu'iei you could say
(u + u') v = Σij (ui + ui') vj( ei fj) = u v + u' v
So I agree, this "construction" using basis functions does produce bilinearity, either in terms of the operator or since t(u+u',v) = t(u,v) + t(u',v) which is to say, in terms of a function t.
I agree that u v and t(u,v) are each uniquely determined by this construction. He again states the idea of bilinear only in that "that is all there is".
In the above paragraph, he does not say that ui are scalars or in some base field or anything like that. Perhaps they are in something like R where multiplication like uivj is at least defined ( the destination space T is "an algebra". )
So the above construction defines t : U x V → T. Since he is saying U V is a tensor product, it must be that the pair (W,t) is a universal pair, and he mentions that, but I see no argument for WHY (W,t) constructed as above would be a universal pair. Where is τ? Where is the proof that τ is unique and linear? [ see below! ] Let's hack a way on the next chunk,
Nothing new in that chunk, so keep going:
OK, he is finally making the claim. I agree with the above equation. But how does that equation show that the function τ is linear?? This is the point I keep missing. I have some kind of misconception right at this point that keeps causing trouble. Here is the triangle from above (with f in place of g)
In my "book" showing that τ is a linear function requires showing that (at least)
τ( eifj + ei'fj') = τ( eifj ) + τ(ei'fj')
At least now I know something about τ (above I knew nothing whatsoever). Again I see that the right side of this equation is the same as
g(ei,fj) + g(ei', fj')
but as usual in all my earlier attempts, I see no way to show that
τ( eifj + ei'fj') = g(ei,fj) + g(ei', fj')
So this is the Kiss of Death in Roman's presentation, for me as a reader. I will try and fail as usual. For example,
τ( eifj + ei'fj') = τ [ t(ei, fj) + t(ei', fj') ] = what is the next step.
The time has now come to pause in Roman until this "misconception" is cleared up. I had the same problem with Ganatra above. I will do some clips from his Problem and Solution on this matter, just to give myself another possible shot that maybe his "solution" really had a solution. φ VW
13.5 Review of Ganatra and Explanation of why τ is linear
Ganatra Problem Statement in HW 7: (this is Problem 3)
Here is a comparison of the Roman and Ganatra triangles,
So Ganatra's student has to show that T is linear (and maybe also unique). Now the homework assignment gives some "hints":
OK, this really is not what I think of as being linear for τ. This is some whole new definition of linear that I do not understand! I will use τ and not T . So his meaning of linear would be, for a simple case
τ( v1w1 + v2w2) = τ(v1w1+ v2w2)
so OK, apart from the two labels being the same, I guess that agrees with my definition of linear. So fine, now the student doing the homework is supposed to use this Hint to show that τ is linear.
The homework assignment has nothing more to say, no more Hints. So I now turn to the homework solution handed out by some TA:
Ganatra Solution Statement for HW 7: (this is Problem 3)
The solution begins like so:
Here again is the Ganatra triangle,
The solver proposes a definition of the function τ
τ (vw) = T(v,w)
I like that, fine. We are then told to "extend this definition linearly to all of VW." So I think τ is linear because it is constructed to be linear. So in his "extension" of the meaning of τ he would say
τ (vw + v'w' ) = T(v,w) + T(v',w')
In other words, τ is linear because we say it is linear. There is nothing to prove.
The triangle picture just shows τ, and we are supposed to construct a τ which makes the triangle commute. The commuting triangle tells us this, for all v and w:
vw = φ(v,w)
x = τ (vw)
x = T(v,w) bilinear
So it tells us that, for all v and w
τ (vw) = T(v,w)
I guess the question is then: if we declare that τ is linear, does that contradict anything in the triangle? The triangle says nothing more than τ (vw) = T(v,w), so forcing τ to be linear does not contradict anything. We are looking for a linear and unique τ, so why not construct it in this manner
(1) τ (vw) = T(v,w)
(2) τ is linear
So I guess we could ask: does this cause any triangle trouble in other ways? Well consider
τ (vw + v'w' ) = τ(v,w) + τ(v',w') // due to our declaration τ is linear
This then says
τ( vw + v'w' ) = x + x'
Now we can write, since T is bilinear
x + x' = T(v,w) + T(v',w')
I cannot combine the T's in this form. But suppose we had a special case that
x' = T(v',w) // no prime on w
Then we would have
x + x' = T(v,w) + T(v',w) = T(v+v',w) = τ ([v+v']w) = τ (φ(v+v',w))
= τ [ φ(v,w) + φ(v',w)] = τ [ φ(v,w)] + τ [ φ(v',w)] = x + x'
The triangle itself tells us that
τ o φ = T
Conclusion: Defining τ to be linear "by fiat" is consistent with the triangle diagram.
So OK, if we do this, why is τ unique? Suppose we had τ and τ' both linear. then
τ o φ = T
τ' o φ = T (τ - τ') o φ = 0
If this is true for ANY φ which defines vw, then I guess τ = τ'. Here is what the solver says:
OK, I am getting happier now. You simply define τ by τ (vw) = T(v,w) according the triangle, and you then declare that τ is linear, and this linear τ then satisfies the triangle and causes no problems, so you can do this declaration as part of the definition of τ. This is what Roman does not make clear.
13.6 Conclusions
1. The tensor product construction which I always make using the basis elements can be associated with a triangle diagram involving two bilinear functions and a linear function τ
There is a trivial way to define τ as τ(uv) = f(u,v), then you "extend" τ by declaring it to be linear, and this is then consistent with the picture. τ is unique as well, and therefore this picture describes (T,t) as a "universal pair". That in turn says that T = U V is a "tensor product"!
2. There is no requirement to do "more". You have a tensor product.
3. The underlying category idea of the triangle is this: the set of bilinear maps t: U x V → T is more or less unique and completely well defined, up to isomorphism, so T can be referred to as U V. For example, in the above triangle picture, if we consider another set of bilinear maps U x V → W where W is a different space from T, there is a 1:1 relationship between maps t: U x V → T and f: U x V → W as embodied in the unique linear function τ which is 1:1, known as a "mediating morphism" (morphism = arrow). In this sense, the set of maps t: U x V → T is "universal" and (T,t) form a "universal pair".
4. You could, however, if you wanted, continue on and define [a b]ij = aibj . This extra add-on fact does not conflict with the triangle or with the definition of tensor product because this add-on declaration is bilinear! That makes it consistent with the earlier stuff.
5. Therefore, an "outer product" is a special case of a "tensor product".
I will try to write up this conclusion more clearly. Ready to re-scan earlier sources maybe!