Section 13

T5

Text-to-text and span corruption

Listen to this chapter

Paper: Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (T5) — Raffel et al., 2020

T5 (Raffel et al.’s 2020 Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer) compared many pre-training choices in a systematic set of experiments. Where GPT and BERT each made one bold bet, T5 ran the systematic study: it changed the objective, architecture, data, and scale in controlled comparisons. These ablationablationAn experiment that removes or changes one part of a system to measure how much that part contributes to the result.See in glossary → studies isolate the contribution of individual choices. Much of what the field “knows” about pre-training design choices traces back to this paper’s experiments.

Everything is text-to-text

T5’s unifying idea is to cast every task (translation, summarization, classification, even regression) as text-to-texttext-to-textT5's framing in which every task — translation, classification, summarization — is cast as "input text → output text", so one model and one objective handle all of them.See in glossary →: feed the model input text, ask it to produce output text. Classification becomes “generate the class name”; translation becomes “generate the translation.” One model, one loss (cross-entropycross-entropy lossA classification loss that penalizes the model according to the negative log-probability it assigned to the correct answer.See in glossary →), one format for everything. This is mostly a fine-tuning/evaluation convenience, but it matters to pre-training because it let T5 compare wildly different tasks on an equal footing, and it foreshadows how today’s models treat all problems as next-token generation.

The text-to-text framework
Translation, classification, similarity scoring, and summarization all become the same job: text in, text out.
Diagram of T5's text-to-text framework: four task inputs (translation, CoLA acceptability, STS-B similarity, summarization) feed into a single T5 model, which emits each answer as text

Figure 1 from Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (Raffel et al., 2020), arXiv 1910.10683.

The span-corruption objective

T5’s pre-training objective generalizes BERT’s masked language modelingmasked language modelMasked Language Model (MLM) — a pre-training objective (used by BERT) that hides a fraction of tokens and trains the model to fill them in using context from both sides. Contrast with next-token prediction.See in glossary → into a sequence-to-sequence denoisingdenoising objectiveAny pre-training objective that corrupts the input (masking, deleting, or shuffling tokens) and trains the model to restore the original. Masked LM and span corruption are both denoising objectives.See in glossary → task called span corruptionspan corruptionT5's pre-training objective: replace random contiguous spans of tokens with sentinel placeholders and train the model to reconstruct the missing spans. A denoising objective.See in glossary →. Instead of masking individual tokens, it masks whole contiguous spans:

  • Randomly select ~15% of tokens; group consecutive selected tokens into spans (mean span length 3).
  • Replace each span in the input with a single unique sentinel token, a marker identifying the missing span, (<X>, <Y>, …).
  • The model’s target is just the dropped spans, each prefixed by its sentinel.

So “the quick brown fox jumps over” might become input “the quick <X> jumps over” with target “<X> brown fox <Y>”. Because one sentinel replaces a whole span, both the corrupted input and the target are short, which makes training efficient compared to predicting every position.

C4: a dataset as a deliverable

To run experiments at scale, T5’s authors built and released the C4C4Colossal Clean Crawled Corpus — the ~750 GB cleaned web-text dataset built from Common Crawl for training T5, and widely reused since.See in glossary → — the Colossal Clean Crawled Corpus — about 750 GB of cleaned English text filtered from Common CrawlCommon CrawlA free, monthly public crawl of the web — petabytes of raw HTML. It is the raw feedstock for most large pre-training corpora after heavy filtering.See in glossary →. The cleaning was aggressive and rule-based (drop pages without terminal punctuation, remove boilerplate and offensive lists, deduplicate). C4 became a standard pre-training dataset in its own right and a template for the heavy filtering pipelines that follow. Releasing the dataset, not just the model, was itself influential.

The architecture, and a note on what won

T5 used a standard encoder-decoderencoder-decoderAn architecture with an encoder that reads the input and a decoder that writes the output, connected by cross-attention. The original transformer and T5 are encoder-decoder models.See in glossary → transformer, and its ablations found that, for their text-to-text setup, encoder-decoder beat decoder-only and encoder-only variants. Scale ran from a 220M-parameter baseline up to an 11-billion-parameter model, trained for 2192^{19} (≈524k) steps on C4 with the memory-efficient AdaFactorAdaFactorA memory-efficient optimizer (used to train T5) that factorizes Adam's second-moment matrix into row and column statistics, drastically cutting optimizer-state memory for very large models.See in glossary → optimizer, an inverse-square-root schedule, and SentencePieceSentencePieceA tokenizer toolkit that operates directly on raw text (treating spaces as symbols), so it works language-agnostically without pre-splitting on whitespace.See in glossary → tokenization (32k vocabulary).

These experiments established several workable architectures and objectives. Comparing much larger training runs also required a way to predict how error would change with model size, data, and computation.