Deep Learning & Neural ArchitecturesLevel: Core MLStudy Battlecard

Transformer Archetypes: Encoder vs Decoder vs Seq2Seq

Structural Differences Between BERT, GPT & T5 for Task-Optimal Architecture Selection

#Architectures#BERT#GPT#T5#Masking#Generative vs Discriminative
Choose Presentation Mode:
STAGE 1 / 7— Anti-Pattern
Section 1: The LLM Anti-Pattern vs Right-Sized Model
The Naive Generative LLM Approach:

Using an autoregressive decoder-only model (GPT) for document classification or dense vector embeddings.

Why It Fails in Production:

Causal decoders only attend to past tokens (left-to-right), throwing away 50% of the bidirectional context that encoders (BERT) leverage naturally.

Targeted Algorithm (Transformer Archetypes: Encoder vs Decoder vs Seq2Seq)
Latency:Encoder: 8ms bidirectional
Cost / 1M Ops:$0.00
Determinism:Full Bidirectional Context
Generative LLM Alternative
Latency:Decoder: 1,500ms causal
Cost / 1M Ops:$5,000
Determinism:Unidirectional Loss of Future Context