Pre-trained model: [Backbone (frozen)] → [New head (trainable)]
Steps:
1. Load pre-trained model (e.g., ResNet-50 on ImageNet)
2. Remove classification head
3. Freeze all backbone parameters
4. Add new task-specific head
5. Train only the head on target data
When: Very small target dataset (< 1000 samples)
Why: Prevents overfitting by keeping pre-trained features intact
Input image → split into patches → mask 75% of patches
│
Visible patches → Encoder (ViT) → latent representations
│
All tokens (visible + mask tokens) → Decoder → reconstruct masked pixels
│
Loss: MSE on masked patches only
Key insight: high masking ratio (75%) forces model to learn semantics,
not just interpolate from nearby patches
Source domain: labeled data (X_s, y_s) ~ P_s(X, Y)
Target domain: unlabeled data X_t ~ P_t(X) (or very few labels)
Goal: train model that performs well on target domain
Challenge: P_s(X) ≠ P_t(X) — different feature distributions
(domain shift / dataset bias / distribution shift)
1. Train model on source labeled data
2. Predict labels for target unlabeled data (pseudo-labels)
3. Filter high-confidence predictions (threshold > 0.95)
4. Retrain model on source data + pseudo-labeled target data
5. Repeat (iteratively improve pseudo-labels)
Variants:
- FixMatch: weak augmentation → pseudo-label, strong augmentation → train
- SHOT: information maximization + self-training (no source data needed)
- AdaMatch: distribution alignment + pseudo-labeling
Support set: K classes, N examples each (K-way N-shot)
Query set: examples to classify
1. Embed all support examples: z = f_θ(x)
2. Compute class prototypes: c_k = (1/N) Σ f_θ(x_i) for class k
3. Classify query by nearest prototype:
P(y=k|x) = softmax(-||f_θ(x) - c_k||²)
Meta-training (learn to learn):
For each task T_i:
1. Inner loop: take few gradient steps on task T_i's support set
θ'_i = θ - α · ∇_θ L(T_i, θ)
2. Outer loop: evaluate adapted model on T_i's query set
Meta-loss += L(T_i, θ'_i)
Update meta-parameters: θ ← θ - β · ∇_θ Σ L(T_i, θ'_i)
At test time:
New task → few gradient steps from θ → adapted model θ'
(fast adaptation because θ is in a region that's easy to adapt from)
Train on Task A → good on A
Train on Task B → good on B, BAD on A (forgotten!)
Train on Task C → good on C, BAD on A and B
Goal: learn tasks sequentially without forgetting previous ones
Objective: find features Φ such that the optimal classifier on top of Φ
is the SAME across all environments (domains)
L_IRM = Σ_e L_e(w ∘ Φ) + λ · ||∇_w L_e(w ∘ Φ)|_{w=1.0}||²
Penalty term: if the gradient at w=1 is non-zero for some environment,
then the representation is not invariant
→ force features that work equally well everywhere
1. Load pre-trained model (e.g., timm.create_model('vit_base', pretrained=True))
2. Replace head for target number of classes
3. Optimizer: AdamW with weight decay 0.01-0.05
4. Learning rate: 1e-4 to 5e-5 (10× less than training from scratch)
5. Schedule: linear warmup (5-10%) + cosine decay
6. Augmentation: RandAugment + Mixup + CutMix
7. Regularization: label smoothing 0.1, stochastic depth
8. Epochs: 10-100 depending on data size
9. EMA (Exponential Moving Average) of weights
Critical: lower lr for backbone, higher lr for head
1. Load pre-trained model (e.g., LLaMA-3-8B)
2. Apply LoRA (rank 8-64, α = 16-32) on q, k, v, o projections
3. Learning rate: 1e-4 to 2e-4
4. Batch size: small (4-16) with gradient accumulation
5. Epochs: 1-5 (less for larger models, more for small datasets)
6. Warmup: 3-10% of steps
7. Context length: pack multiple examples per sequence
8. Data: high-quality instruction-response pairs
9. Evaluation: held-out set + task-specific benchmarks