Skip to main content

Transformer Architecture · Chapter 05

Multi-Head Attention

h parallel heads, same cost, richer representation

This chapter is part of the Pro track. Pro beta members receive the complete self-paced architecture sequence.

Related reading · Knowledge Lab

Build the context for this chapter