ROC (Receiver Operating Characteristic): Plot TPR vs FPR at all thresholds.
TPR (True Positive Rate) = Recall = TP / (TP+FN) (y-axis)
FPR (False Positive Rate) = FP / (FP+TN) (x-axis)
AUC = Area Under ROC Curve
= 0.5 → random classifier
= 1.0 → perfect classifier
= Probability that model ranks a random positive higher than a random negative
When to use ROC-AUC:
- Balanced datasets
- Equal cost of FP and FN
- Need threshold-independent comparison
When NOT to use ROC-AUC:
- Highly imbalanced data (FPR stays low even with many FP because TN dominates)
- When you care more about precision than overall discrimination
Precision (y-axis) vs Recall (x-axis) at all thresholds
AUPRC (Average Precision):
AP = Σ_n (R_n - R_{n-1}) · P_n
≈ Area under the PR curve
When to use AUPRC:
- Imbalanced datasets (positive class is rare)
- Cost of false positives and false negatives differ significantly
- Fraud, medical diagnosis, anomaly detection
R² = 1 - Σ(yᵢ - ŷᵢ)² / Σ(yᵢ - ȳ)²
R² = 1.0 → perfect prediction
R² = 0.0 → model is as good as predicting the mean
R² < 0.0 → model is WORSE than predicting the mean (very bad)
DCG@K = Σ_{k=1}^{K} (2^{rel_k} - 1) / log₂(k + 1)
IDCG@K = DCG of ideal ranking (sorted by relevance)
NDCG@K = DCG@K / IDCG@K
rel_k = relevance score of item at position k (can be graded: 0, 1, 2, 3)
NDCG properties:
- Handles graded relevance (not just binary)
- Position-aware (higher ranks matter more)
- Normalized to [0, 1] — comparable across queries
- Standard for search and recommendation systems
For each sample i:
a(i) = mean distance to other points in same cluster (cohesion)
b(i) = mean distance to points in nearest other cluster (separation)
s(i) = (b(i) - a(i)) / max(a(i), b(i))
Overall silhouette = mean(s(i) for all i)
s = +1: point well-clustered
s = 0: point on border between clusters
s = -1: point likely in wrong cluster
Fold 1: Train [────────] Val [──]
Fold 2: Train [──────────────] Val [──]
Fold 3: Train [────────────────────] Val [──]
Fold 4: Train [──────────────────────────] Val [──]
Key: training always BEFORE validation (no future leakage)
Outer loop (evaluation — unbiased performance estimate):
for each outer fold:
Inner loop (tuning — find best hyperparameters):
for each inner fold:
train with candidate hyperparams
evaluate on inner val
select best hyperparams
Train with best hyperparams on full outer train
Evaluate on outer test → final performance estimate
Result: unbiased estimate of generalization performance
for b in range(B): # B = 1000-10000
sample = resample(test_predictions) # sample with replacement
metric_b = compute_metric(sample)
CI_lower = percentile(metrics, 2.5)
CI_upper = percentile(metrics, 97.5)
If CIs of two models don't overlap → statistically significant difference
A model is calibrated if when it predicts 70% probability, the event occurs ~70% of the time.
Perfectly calibrated:
P(y=1 | predicted_prob = 0.7) = 0.7
Over-confident: model says 0.9, actual rate is 0.6
Under-confident: model says 0.5, actual rate is 0.8