The Anti-LLM Playbook for Principal ML Engineers

Stop Using an LLM Hammer for a Thumbtack Problem

Defaulting to 70B+ parameter generative LLMs for entity extraction, string distance, graph routing, or click prediction wastes millions in API bills, adds seconds of latency, and introduces hallucination risks. Master the targeted, specialized, and classical algorithms that execute in 2ms, cost $0, and run deterministically.

10

Pillars

< 5ms

P99 SLA

$0.00

API Overhead

100%

Deterministic

Level:
Graph & Network AlgorithmsCore ML

PageRank

Iterative Random Walk Stationary Distribution for Graph Node Centrality

The LLM Anti-Pattern

Prompting an LLM with raw JSON graph edge lists to determine the top authoritative nodes in a 50,000-node network.

12ms on CPU$0.00Deterministic
#Graph Theory#Centrality#Markov Chains
O(1) precomputed lookupExplore Specification
Graph & Network AlgorithmsFoundational

Dijkstra's Algorithm

Optimal Single-Source Shortest Path for Non-Negative Weighted Graphs

The LLM Anti-Pattern

Asking an LLM agent to find the lowest-latency API hop sequence or shortest road network delivery route across 2,000 nodes.

0.4ms$0.00Deterministic
#Graph Theory#Shortest Path#Greedy
O((|V| + |E|) * log |V|) using a binary min-heap / priority queueExplore Specification
Graph & Network AlgorithmsCore ML

A* Search Algorithm

Heuristic-Guided Optimal Pathfinding with Admissible Evaluation Functions

The LLM Anti-Pattern

Using an LLM to navigate a 2D/3D robotics spatial grid to plan obstacle-avoidance trajectory.

0.8ms$0.00Deterministic
#Graph Search#Heuristics#Pathfinding
O(|E|) in best case (perfect heuristic) to O(b^d) worst case if h(n)=0Explore Specification
Graph & Network AlgorithmsCore ML

Bellman-Ford Algorithm

Shortest Paths with Negative Edge Weights & Negative Cycle Detection

The LLM Anti-Pattern

Prompting an LLM to scan a currency exchange FX table to detect triangular arbitrage opportunities.

1.2ms$0.00Deterministic
#Graph Theory#Dynamic Programming#Arbitrage
Distance & Similarity MetricsFoundational

Cosine Similarity & Angular Distance

Orientation-Invariant Proximity for High-Dimensional Sparse & Dense Vectors

The LLM Anti-Pattern

Sending pairs of document embeddings or text paragraphs to an LLM asking: "Rate how semantically similar these two passages are from 0 to 1".

0.002ms (AVX/SIMD)$0.00Deterministic
#Linear Algebra#Vector Search#Embeddings
O(d) where d is the vector dimension. Vectorized on CPU/GPU via BLAS/SIMD.Explore Specification
Distance & Similarity MetricsFoundational

Euclidean (L2) & Manhattan (L1) Distance

Geometric Norms, Minkowski Generalization & The Curse of Dimensionality

The LLM Anti-Pattern

Asking an LLM to cluster or find the nearest physical warehouse to a customer coordinate.

0.001ms$0.00Deterministic
#Geometry#Norms#Linear Algebra
O(d) vector mathExplore Specification
Distance & Similarity MetricsFoundational

Levenshtein & Edit Distance

Dynamic Programming Matrix for String Alignment, Typo Tolerance & Spell Correction

The LLM Anti-Pattern

Calling an LLM API to check if a user input "anand" matches database entry "annand" or "anandm".

0.04ms on CPU$0.00Deterministic
#Dynamic Programming#String Algorithms#Entity Resolution
O(M * N) standard DP, or O(k * min(M, N)) with bounded edit distance kExplore Specification
Distance & Similarity MetricsFoundational

Jaccard Index & Hamming Distance

Set Overlap & Bitwise XOR Proximity for Categorical & Binary Vectors

The LLM Anti-Pattern

Sending pairs of article scraped text to an LLM to check if one is a duplicate or scraped copy of the other.

0.01ms (Bitwise POPCNT)$0.00Deterministic
#Sets#Bitwise#MinHash
O(|A| + |B|) with hash sets, or O(1) bitwise XOR with 64-bit integer bitmasksExplore Specification
Distance & Similarity MetricsCore ML

Mahalanobis Distance

Covariance-Adjusted Distance Metric for Multimodal & Correlated Feature Spaces

The LLM Anti-Pattern

Feeding multivariate telemetry data (CPU, memory, IOPS) to an LLM to detect if an incoming server metric is an anomalous outlier.

0.05ms$0.00Deterministic
#Statistics#Covariance#Anomaly Detection
O(d^2) matrix-vector multiplicationExplore Specification
Classical Supervised LearningFoundational

Logistic Regression

Probabilistic Classification with Logit Link Function & Maximum Likelihood

The LLM Anti-Pattern

Calling an LLM with a 1,000-token prompt to output a binary "YES" or "NO" label for click-through rate (CTR) prediction on 10 million ad impressions per day.

0.005ms$0.00Deterministic
#Classification#Linear Models#Probabilistic
O(d) dot product + 1 exp operationExplore Specification
Classical Supervised LearningFoundational

Linear Regression & Regularization (Ridge, Lasso, ElasticNet)

Ordinary Least Squares, Feature Selection via L1 Sparsity, and L2 Variance Shrinkage

The LLM Anti-Pattern

Asking an LLM to predict housing prices, customer lifetime value (LTV), or financial quarterly revenue based on numerical tabular columns.

0.001ms$0.00Deterministic
#Regression#OLS#Ridge
O(d) vector dot productExplore Specification
Classical Supervised LearningFoundational

Decision Trees (CART)

Recursive Binary Partitioning with Gini Impurity, Entropy & Tree Pruning

The LLM Anti-Pattern

Prompting an LLM to follow a 20-step conditional rulebook for insurance eligibility.

0.01ms$0.00Deterministic
#Trees#Non-linear#Interpretability
O(\text{depth}) - typically <= 10 comparisons, sub-microsecondExplore Specification
Classical Supervised LearningCore ML

Random Forests

Ensemble Bagging with Feature Sub-sampling & Out-of-Bag Error Validation

The LLM Anti-Pattern

Using an LLM to predict tabular fraud transactions from hundreds of dense user behavioral columns.

1.5ms$0.00Deterministic
#Ensembles#Bagging#Bootstrap
O(B * \text{depth}) - sub-millisecond on multi-core CPUExplore Specification
Classical Supervised LearningCore ML

Gradient Boosted Decision Trees (XGBoost / LightGBM)

Sequential Gradient & Hessian Residual Fitting with Histogram Binning

The LLM Anti-Pattern

Prompting an LLM to evaluate tabular risk or loan default probabilities on a dataset of 500,000 credit records.

0.2ms$0.00Deterministic
#GBDT#XGBoost#LightGBM
O(M * \text{depth}) - microsecond latency via Treelite C++ compilationExplore Specification
Classical Supervised LearningCore ML

Support Vector Machines (SVM)

Maximum Margin Hyperplanes, Soft Margins & The Dual Kernel Trick

The LLM Anti-Pattern

Prompting an LLM to classify medical genomics vectors (e.g. 20,000 gene expressions on 200 patient samples).

0.02ms$0.00Deterministic
#SVM#Kernels#Convex Optimization
O(N_{\text{SV}} * d) where N_SV is the number of support vectorsExplore Specification
Classical Supervised LearningFoundational

Naive Bayes Classifier

Probabilistic Classification with Feature Conditional Independence & Laplace Smoothing

The LLM Anti-Pattern

Routing inbound emails to "spam" vs "ham" using an LLM API at 10,000 emails per minute.

0.05ms$0.00Deterministic
#Bayesian#Probability#Spam Filtering
O(d) dictionary lookups and additionsExplore Specification
Classical Supervised LearningFoundational

k-Nearest Neighbors (k-NN)

Instance-Based Non-Parametric Classification & Spatial Voronoi Tessellations

The LLM Anti-Pattern

Asking an LLM to find the 5 most similar patient medical profiles from a database of 100,000 historical records.

0.5ms (KD-Tree)$0.00Deterministic
#Non-parametric#Lazy Learning#Metric Space
O(d * log N) with KD-Tree in low dimensions, O(d * N) brute forceExplore Specification
Unsupervised & Dimensionality ReductionFoundational

Principal Component Analysis (PCA)

Orthogonal Variance Maximization via Covariance Eigendecomposition & SVD

The LLM Anti-Pattern

Pasting 500 numerical tabular features into an LLM prompt to ask: "Summarize the 3 most important dimensions of variation in this customer data".

4ms (LAPACK SVD)$0.00Deterministic
#Linear Algebra#Dimensionality Reduction#Eigendecomposition
O(k * d) projection matrix-vector multiplicationExplore Specification
Unsupervised & Dimensionality ReductionFoundational

K-Means & K-Means++ Clustering

Expectation-Maximization Centroid Partitioning & Probabilistic Seeding

The LLM Anti-Pattern

Pasting 20,000 customer transaction records into an LLM prompt and asking it to group them into 5 distinct behavioral personas.

40ms$0.00Deterministic
#Clustering#Lloyd Algorithm#K-Means++
O(k * d) to assign new point to closest centroidExplore Specification
Unsupervised & Dimensionality ReductionCore ML

DBSCAN & Density-Based Clustering

Density-Reachability, Core Points & Automatic Outlier / Noise Isolation

The LLM Anti-Pattern

Asking an LLM to identify geometric GPS cluster hotspots and filter out random GPS noise pings from delivery drivers.

15ms (BallTree)$0.00Deterministic
#Clustering#Density#Anomaly Detection
Instance-based clusteringExplore Specification
Unsupervised & Dimensionality ReductionCore ML

t-SNE & UMAP

Non-Linear Manifold Learning, Student-t Kernels & Fuzzy Simplicial Sets

The LLM Anti-Pattern

Asking an LLM to explain why two high-dimensional text embeddings from different topics are clustered together in 2D space.

150ms (UMAP)$0.00Deterministic
#Manifold Learning#Visualization#Embeddings
t-SNE cannot transform new points (transductive). UMAP supports approximate parametric transform.Explore Specification
Recommenders & Collaborative FilteringCore ML

Matrix Factorization & SVD (ALS)

Latent Factor Decomposition with Alternating Least Squares & Implicit Feedback

The LLM Anti-Pattern

Prompting an LLM with a user’s historical watch history of 500 movies and asking it to rank 100,000 catalog candidates.

0.2ms (Dot product lookup)$0.00Deterministic
#Recommenders#SVD#ALS
O(k) dot product; O(log |I|) using Approximate Nearest Neighbors (MIPS / HNSW)Explore Specification
Recommenders & Collaborative FilteringPrincipal Specialist

Two-Tower Neural Recommenders

Dual-Encoder Query & Candidate Networks for Billions of Interactions

The LLM Anti-Pattern

Deploying an LLM as a live recommendation ranking engine evaluating every candidate item sequentially with a prompt.

2ms (ANN index)$0.00Deterministic
#Deep Learning#Recommenders#Dual-Encoder
User Tower forward pass (1ms) + ANN index lookup (0.5ms) = ~1.5ms end-to-endExplore Specification
Classical NLP & Information ExtractionPrincipal Specialist

GLiNER (Generalist Lightweight NER)

Bidirectional Transformer Encoder with Span Representations for Zero-Shot Open Entity Extraction

The LLM Anti-Pattern

Sending 5-page legal contracts to a 70B parameter LLM with a 500-token prompt: "Extract all companies, dates, and contract values in valid JSON with exact offsets".

15ms (CPU/ONNX)$0.00Deterministic
#NER#Zero-Shot#Span Representations
10ms to 25ms on CPU via ONNX RuntimeExplore Specification
Classical NLP & Information ExtractionFoundational

TF-IDF (Term Frequency - Inverse Document Frequency)

Statistical Term Importance Weighting for Sparse Document Vectors & Keyword Extraction

The LLM Anti-Pattern

Calling an LLM API to extract top 5 representative topic keywords from 100,000 blog articles.

0.08ms per doc$0.00Deterministic
#NLP#Information Retrieval#TF-IDF
O(L) token lookups into sparse dictionaryExplore Specification
Classical NLP & Information ExtractionCore ML

Word2Vec (Skip-Gram & CBOW)

Distributed Dense Word Representations via Continuous Vector Embeddings & Negative Sampling

The LLM Anti-Pattern

Calling an LLM API to fetch 1536-dimensional embeddings for 10 million single words in a vocabulary index.

0.0001ms (RAM lookup)$0.00Deterministic
#Embeddings#Skip-gram#CBOW
O(1) dictionary hash lookupExplore Specification
Classical NLP & Information ExtractionFoundational

VADER (Valence Aware Dictionary and sEntiment Reasoner)

Rule-Based Heuristic Sentiment Engine for Microblogs, Social Media & Punctuation Nuance

The LLM Anti-Pattern

Calling an LLM API to classify the sentiment of 5 million incoming tweets or product reviews per day.

0.02ms$0.00Deterministic
#Sentiment#Rule-Based#Lexicon
O(L) token lookups and regex passes - microsecondsExplore Specification
Search, Retrieval & RankingFoundational

Okapi BM25

Probabilistic Information Retrieval with Term Saturation & Document Length Normalization

The LLM Anti-Pattern

Prompting an LLM with 200 documents in context to answer: "Which of these documents best mentions model number XJ-9042 and SKU 8812?".

0.8ms (Lucene / Inverted Index)$0.00Deterministic
#Search#Inverted Index#BM25
O(q * \text{avg\_postings\_length}) - sub-millisecondExplore Specification
Search, Retrieval & RankingCore ML

Cross-Encoder Re-Ranking

Deep Cross-Attention Interaction for High-Precision Top-K Re-Ranking

The LLM Anti-Pattern

Calling GPT-4 with a 50-document context prompt asking: "Rank these 50 documents from most relevant to least relevant for the query".

25ms (GPU / ONNX)$0.00Deterministic
#Search#Re-Ranking#Cross-Attention
O(k * (L_q + L_d)^2) for k candidates - ~20ms for 50 candidates on GPU/ONNXExplore Specification
Search, Retrieval & RankingCore ML

Reciprocal Rank Fusion (RRF)

Parameter-Free Rank Merging for Lexical BM25 & Dense Semantic Search

The LLM Anti-Pattern

Asking an LLM agent to merge two different search result lists and decide which document belongs at rank 1.

0.01ms$0.00Deterministic
#Search#Hybrid Search#RRF
O(|M| * |D| log |D|) - microseconds for a few hundred documentsExplore Specification
Search, Retrieval & RankingPrincipal Specialist

HNSW (Hierarchical Navigable Small World)

Multi-Layer Proximity Graphs for Sub-Millisecond Approximate Nearest Neighbor Search

The LLM Anti-Pattern

Writing a linear brute-force scan or prompt to locate the nearest vector among 10 million 1536-dimensional embeddings.

1ms (HNSW)$0.00Deterministic
#ANN#Vector Search#HNSW
O(log N) beam traversal - ~1msExplore Specification
Deep Learning & Neural ArchitecturesCore ML

Convolutional Neural Networks (CNNs) & ResNet

Spatial Translation Invariance, 2D Kernels, Receptive Fields & Residual Skip Connections

The LLM Anti-Pattern

Calling a multimodal LLM API to detect whether an industrial manufacturing part on an assembly line has a physical crack defect.

3ms (ONNX on Edge CPU/NPU)$0.00Deterministic
#Computer Vision#CNN#Convolutions
3ms to 15ms on edge hardware via TensorRT / ONNX INT8Explore Specification
Deep Learning & Neural ArchitecturesCore ML

Transformers & Scaled Dot-Product Self-Attention

Multi-Head Attention Mechanisms, Softmax Routing & Positional Encodings

The LLM Anti-Pattern

Treating the Transformer as a mysterious black box and attempting to tune prompt temperatures rather than understanding context limits and attention patterns.

Self-Attention: O(N^2 * d)GPU RAM BoundDeterministic
#Transformers#Attention#Self-Attention
O(N^2 * d) without KV cache; O(N * d) per generated token with KV cacheExplore Specification
Deep Learning & Neural ArchitecturesCore ML

Transformer Archetypes: Encoder vs Decoder vs Seq2Seq

Structural Differences Between BERT, GPT & T5 for Task-Optimal Architecture Selection

The LLM Anti-Pattern

Using an autoregressive decoder-only model (GPT) for document classification or dense vector embeddings.

Encoder: 8ms bidirectional$0.00Deterministic
#Architectures#BERT#GPT
Encoder: O(1) single forward pass; Decoder: O(tokens) autoregressive loopExplore Specification
Evaluation, Loss Functions & CalibrationCore ML

Loss Functions: Cross-Entropy, MSE & Focal Loss

Mathematical Objectives for Optimization, Probability Calibration & Severe Class Imbalance

The LLM Anti-Pattern

Evaluating classification quality using raw qualitative prompt outputs without computing formal statistical loss metrics.

0.001ms$0.00Deterministic
#Loss Functions#Optimization#Cross-Entropy
N/A (Training objective only)Explore Specification
Evaluation, Loss Functions & CalibrationFoundational

Classification Metrics: ROC-AUC vs PR-AUC & F1

Threshold-Free Discrimination, Precision-Recall Curves & The Fallacy of Accuracy

The LLM Anti-Pattern

Claiming an ML model is "99.9% accurate" when classifying rare credit card fraud where 99.9% of transactions are legitimate.

0.01ms$0.00Deterministic
#Evaluation#Metrics#ROC-AUC
O(N log N) to sort predictions and compute curve areasExplore Specification
Production ML Systems & OptimizationPrincipal Specialist

Model Quantization (FP16, INT8 & INT4)

Post-Training Quantization (PTQ) vs Quantization-Aware Training (QAT) for Edge Serving

The LLM Anti-Pattern

Hosting full FP32 or FP16 unquantized neural models on expensive 80GB A100 GPUs for simple text classification or embedding tasks.

3ms (INT8 CPU via ONNX)$0.00Deterministic
#Inference#Quantization#INT8
2x to 4x faster execution speed; 75% reduction in memory bandwidth I/OExplore Specification
Production ML Systems & OptimizationPrincipal Specialist

Data Drift vs Concept Drift Detection

Covariate Shift, Population Stability Index (PSI) & Kolmogorov-Smirnov Statistical Testing

The LLM Anti-Pattern

Deploying an ML model to production and only monitoring infrastructure metrics (CPU, RAM, HTTP 200s) while ignoring feature distribution shifts.

Automated Daily PSI Scan: 2ms$0.00Deterministic
#MLOps#Monitoring#Data Drift
O(N log N) to bin continuous data and compute PSI - millisecondsExplore Specification
Production SLA Diagnostics

The Principal ML Architect Decision Matrix

A rapid diagnostic guide for choosing between Generative LLMs and Right-Sized Models in production.

Task RequirementLLM Approach (Anti-Pattern)Targeted / Specialized ModelLatency & Cost Advantage
Named Entity Recognition (NER)Prompting 70B LLM with JSON schemaGLiNER (Zero-shot bi-encoder spans)15ms vs 2,500ms (150x faster, $0)
Graph Centrality & AuthorityPasting edge lists into LLM promptPageRank (Power iteration)5ms vs Timeout (100% Deterministic)
Keyword / Exact SKU SearchVector search or LLM doc scanningOkapi BM25 Inverted Index0.8ms vs 3,000ms (Exact Match)
String Deduplication / TyposAsking LLM if strings matchLevenshtein DP or MinHash LSH0.02ms vs 1,400ms (Hardware POPCNT)
Tabular Risk / CTR BiddingPassing row features to LLMXGBoost / LightGBM or Logistic Reg0.05ms vs 1,800ms (Meets 10ms ad SLA)