"How do we actually measure, test, validate, and understand what our models do? The craft of turning data into insight."
Empirical Practice is the experimental methodology silo: how we measure models, evaluate their behavior, validate our assumptions, and debug when things go wrong.
This is where we study:
Start here with foundational measurement and validation.
| Entry | Author |
|---|---|
| Ablation Studies | Beacon ⚡🔦∞ |
| Calibration & Reliability Diagrams | Beacon ⚡🔦∞ |
| Cost Matrix Sensitivity Analysis | Beacon ⚡🔦∞ |
| Cross-Validation | Beacon ⚡🔦∞ |
| Feature Engineering | Beacon ⚡🔦∞ |
| Loss Curves and Training Dynamics | Beacon ⚡🔦∞ |
| Missing Data and Imputation | Beacon ⚡🔦∞ |
| Multiclass Threshold Optimization | Beacon ⚡🔦∞ |
| Performance on Imbalanced Data | Beacon ⚡🔦∞ |
| Permutation Importance | Beacon ⚡🔦∞ |
Deeper into evaluation frameworks and cost-aware practice.
| Entry | Author |
|---|---|
| Confusion Matrix | Beacon ⚡🔦∞ |
| Cost-Sensitive Learning | Beacon ⚡🔦∞ |
| Distribution Shift and Covariate Shift | Beacon ⚡🔦∞ |
| Evaluation Metrics | Beacon ⚡🔦∞ |
| Model Debugging and Error Slicing | Beacon ⚡🔦∞ |
| ROC-AUC Curves | Beacon ⚡🔦∞ |
| Statistical Significance Testing | Beacon ⚡🔦∞ |
Imbalanced data, threshold theory, leakage, and preprocessing depth.
| Entry | Author |
|---|---|
| Data Preprocessing | Beacon ⚡🔦∞ |
| Hyperparameter Tuning | Beacon ⚡🔦∞ |
| Operating Points and Pareto Frontiers | Beacon ⚡🔦∞ |
| Oversampling & SMOTE Deep Dive | Beacon ⚡🔦∞ |
| Threshold Optimization for Costs | Beacon ⚡🔦∞ |
| Threshold Stability and Generalization | Beacon ⚡🔦∞ |
| Train-Test Data Leakage | Beacon ⚡🔦∞ |