Principal ML Engineer Architectural Cheat-Sheet
Condensed mathematical formulas, Big-O complexities, and anti-LLM model selection rules.
PragmaticML — Principal ML Architectural Cheat-Sheet
Anand Muraleedharan • anandmuraleedharan.com • Right-Sized ML Reference
1. Mathematical Metric & Distance Formulations
Length-invariant vector orientation. Optimal for dense embeddings.
L1 is robust to outliers; L2 penalizes large deviations quadratically.
Covariance-adjusted. Invariant to scale and feature correlations.
Set overlap ratio. Foundation of MinHash Locality-Sensitive Hashing.
Sub-linear term frequency saturation with document length penalty.
Stationary Markov chain distribution with damping d=0.85.
2. The Right-Size Model Selection Matrix
| Engineering Task | Right-Sized Algorithm | Why Not An LLM? | Inference SLA |
|---|---|---|---|
| Named Entity Extraction (NER) | GLiNER (DeBERTa Spans) | Hallucinates keys; fails character offsets | 15ms (CPU) |
| Exact SKU / Codebase Search | Okapi BM25 Inverted Index | Lost in the middle; token subword blur | 0.8ms (CPU) |
| RAG Top-50 Re-Ranking | Cross-Encoder (MiniLM / BGE) | Position bias; costs $0.05 per prompt | 20ms (ONNX) |
| Real-Time CTR / Loan Risk | Logistic Regression / LightGBM | Violates 10ms SLA; uncalibrated outputs | 0.05ms (C++) |
| Graph Routing / Shortest Path | Dijkstra / A* Search | Hallucinates edges; non-optimal hops | 0.4ms (Heap) |
| Streaming Sentiment (50k/sec) | VADER Rule Engine | Costs $10k/day; network jitter | 0.02ms (RAM) |
| 10M Catalog Recommendations | Two-Tower Encoders + HNSW | Cannot evaluate 10M dot products | 2ms (MIPS) |
3. Algorithmic Time & Space Complexity Reference
4. Optimization Loss Objectives
Proper scoring rule. Outputs are calibrated probabilities.
Down-weights easy negatives. Requires post-hoc Platt scaling.
Scale-invariant rank merger for Hybrid Search.
1. Never use an autoregressive LLM when a bidirectional encoder (BERT / GLiNER) suffices. • 2. For tabular data, GBDTs (XGBoost/LightGBM) outperform deep neural nets in 95% of production benchmarks. • 3. Always pair dense vector search with sparse BM25 to prevent catastrophic recall failure on exact part numbers and error codes. • 4. Quantize linear layers to INT8 via PTQ before provisioning expensive GPU cloud instances.