Smol and Ultra-Scale Training

view markdown

Session One: An Introduction to Training and Memory-Overheads

Session Three: An Overview of Attention in the Transformer