How repeating the same Transformer layers can help, why the model needs training for it, and what extra passes cost in time and memory.