Skip to main content

Transformer Architecture · Chapter 02

Token Embeddings

One-hot × Wₑ × √d_model — sparse IDs to dense vectors

This chapter is part of the Pro track. Pro beta members receive the complete self-paced architecture sequence.

Related reading · Knowledge Lab

Build the context for this chapter