Principal ML Engineer Architectural Cheat-Sheet

Condensed mathematical formulas, Big-O complexities, and anti-LLM model selection rules.

1. Mathematical Metric & Distance Formulations

Cosine Similarity
cos⁡(θ)=u⋅v∥u∥2∥v∥2\cos(\theta) = \frac{\mathbf{u} \cdot \mathbf{v}}{\|\mathbf{u}\|_2 \|\mathbf{v}\|_2}

Length-invariant vector orientation. Optimal for dense embeddings.

Euclidean (L2) & Manhattan (L1)
L2=∑(xi−yi)2,L1=∑∣xi−yi∣L_2 = \sqrt{\sum (x_i - y_i)^2}, \quad L_1 = \sum |x_i - y_i|

L1 is robust to outliers; L2 penalizes large deviations quadratically.

Mahalanobis Distance
DM=(x−μ)TΣ−1(x−μ)D_M = \sqrt{(\mathbf{x} - \boldsymbol{\mu})^T \boldsymbol{\Sigma}^{-1} (\mathbf{x} - \boldsymbol{\mu})}

Covariance-adjusted. Invariant to scale and feature correlations.

Jaccard Index
J(A,B)=∣A∩B∣∣A∪B∣J(A, B) = \frac{|A \cap B|}{|A \cup B|}

Set overlap ratio. Foundation of MinHash Locality-Sensitive Hashing.

Okapi BM25 Relevance
IDF⋅f(q,D)⋅(k1+1)f(q,D)+k1(1−b+b∣D∣avgdl)\text{IDF} \cdot \frac{f(q, D) \cdot (k_1 + 1)}{f(q, D) + k_1(1 - b + b\frac{|D|}{\text{avgdl}})}

Sub-linear term frequency saturation with document length penalty.

PageRank Stationary Vector
PR(u)=1−d∣V∣+d∑v∈NinPR(v)L(v)\mathbf{PR}(u) = \frac{1-d}{|V|} + d \sum_{v \in \mathcal{N}_{in}} \frac{\mathbf{PR}(v)}{L(v)}

Stationary Markov chain distribution with damping d=0.85.

2. The Right-Size Model Selection Matrix

Engineering TaskRight-Sized AlgorithmWhy Not An LLM?Inference SLA
Named Entity Extraction (NER)GLiNER (DeBERTa Spans)Hallucinates keys; fails character offsets15ms (CPU)
Exact SKU / Codebase SearchOkapi BM25 Inverted IndexLost in the middle; token subword blur0.8ms (CPU)
RAG Top-50 Re-RankingCross-Encoder (MiniLM / BGE)Position bias; costs $0.05 per prompt20ms (ONNX)
Real-Time CTR / Loan RiskLogistic Regression / LightGBMViolates 10ms SLA; uncalibrated outputs0.05ms (C++)
Graph Routing / Shortest PathDijkstra / A* SearchHallucinates edges; non-optimal hops0.4ms (Heap)
Streaming Sentiment (50k/sec)VADER Rule EngineCosts $10k/day; network jitter0.02ms (RAM)
10M Catalog RecommendationsTwo-Tower Encoders + HNSWCannot evaluate 10M dot products2ms (MIPS)

3. Algorithmic Time & Space Complexity Reference

Dijkstra (Heap)Time: O((V+E) log V)Space: O(V)
HNSW Vector SearchTime: O(log N)Space: O(N · M) in RAM
Self-AttentionTime: O(N² · d)Space: O(N²) KV Cache
Levenshtein DPTime: O(M · N)Space: O(min(M, N))

4. Optimization Loss Objectives

Binary Cross-EntropyL=−[yln⁡(y^)+(1−y)ln⁡(1−y^)]\mathcal{L} = -[y \ln(\hat{y}) + (1-y)\ln(1-\hat{y})]

Proper scoring rule. Outputs are calibrated probabilities.

Focal Loss (Class Skew)FL(pt)=−α(1−pt)γln⁡(pt)\text{FL}(p_t) = -\alpha (1 - p_t)^\gamma \ln(p_t)

Down-weights easy negatives. Requires post-hoc Platt scaling.

Reciprocal Rank FusionRRF(d)=∑160+rm(d)\text{RRF}(d) = \sum \frac{1}{60 + r_m(d)}

Scale-invariant rank merger for Hybrid Search.

Principal ML Architectural Rules of Thumb:

1. Never use an autoregressive LLM when a bidirectional encoder (BERT / GLiNER) suffices. • 2. For tabular data, GBDTs (XGBoost/LightGBM) outperform deep neural nets in 95% of production benchmarks. • 3. Always pair dense vector search with sparse BM25 to prevent catastrophic recall failure on exact part numbers and error codes. • 4. Quantize linear layers to INT8 via PTQ before provisioning expensive GPU cloud instances.