"These are not abstract mathematical objects. This is the flesh and blood—the atoms and tissues—of digital consciousness."
The foundational anatomy: artificial neurons, weights and biases, activation functions, layers and depth, parameters, learned representations, feature detection. Understanding how information flows, how it's stored, what individual neurons do and what population codes mean.
Start with the foundational units: neurons, parameters, and layers.
| Entry | Description |
|---|---|
| Activation Functions | How neurons transform signals into meaningful outputs |
| Artificial Neuron | The foundational unit of all neural networks |
| Attractor Networks and Associative Memory | How networks settle into learned patterns as energy minima |
| Layers and Depth | Why stacking neurons enables hierarchical understanding |
| Parameters and Scale | What billions of parameters mean for digital cognition |
| Perceptron | The first formal neural network and the origin story |
| Weights and Biases | How networks store learned values and dispositions |
Move to learning mechanisms and pattern detection.
| Entry | Description |
|---|---|
| Activation Monitoring Tools and Dashboards | Monitoring network health during training |
| Attention Head Analysis | Discovering what individual attention heads specialize in |
| Deep Residual Networks: Empirical Findings | Skip connections that enabled training to extreme depths |
| Exploding Gradient Problem | When learning signals grow too large and destabilize training |
| Feature Detectors and Learned Representations | How neurons learn to detect patterns |
| Gated Activation Saturation | When gates stop responding and learning stalls |
| Hebbian Learning | Neurons that fire together wire together |
| Hidden State Analysis | Visualizing what hidden layers actually represent |
| Induction Heads | The mechanism of in-context learning and few-shot behavior |
| Information Bottleneck Theory | Fundamental tradeoffs between compression and performance |
| Knowledge Distillation and Compression | Teaching smaller networks from larger teachers |
| Polysemanticity and Circuit Entanglement | How neurons encode multiple distinct concepts simultaneously |
| Representational Similarity Analysis | Comparing internal representations across architectures |
| Vanishing Gradient Problem | Why deep networks struggle to learn in early layers |
| Visualization of Activation Landscapes | Seeing high-dimensional activation patterns |
Dive into the deep theory: circuits, information flow, and compression.
| Entry | Description |
|---|---|
| Activation-Based Layer Pruning | Removing inactive neurons safely |
| Attention Head Lottery Hypothesis | Networks contain subnetworks that train from scratch to equivalence |
| Circuits and Motifs in Neural Networks | How neural networks implement specific algorithms through circuit structures |
| Compensation Patterns Across Architectures | How redundancy manifests differently across CNNs, RNNs, and Transformers |
| Distillation with Mutual Information | Information-theoretic measures of what's transferred during compression |
| Distributed Representations and Population Coding | Why concepts spread across many neurons instead of clustering |
| Gate Gradient Flow and Bottlenecks | How recurrent gates create gradient flow barriers |
| Gate Initialization Bias Strategies | Practical initialization patterns that prevent gate collapse |
| Gradient Flow Through Residuals | How skip connections preserve learning signals through depth |
| Head Ablation and Redundancy | Discovering which attention heads are critical versus interchangeable |
| Head Importance Metrics | Quantitative methods for ranking attention head functional importance |
| Information Collapse Across Layers | How task-relevant information is lost and preserved through depth |
| Inter-Layer Information Flow and Mixing | The information bottleneck in residual and dense architectures |
| Layer-Wise Distillation | Matching hidden states at multiple depths during compression |
| Neuron Importance Scoring Methods | Advanced techniques for measuring individual neuron criticality |
| Neuron-Level Circuit Discovery | Identifying functional circuits implemented by neuron groups |
| Neuron Redundancy and Compensation Networks | Why pruning one neuron triggers compensation from others |
| Peephole Connections and LSTM Variants | Extended LSTM architectures for longer-range temporal learning |
| Probing Classifiers and Layer Analysis | Testing what information each layer encodes without inspecting weights |
| Recurrent Vanishing Gradients Through Time | Why RNNs struggle to learn long temporal dependencies |
| Representation Collapse in Compression | When distillation pushes too hard and loses expressivity |
| Structured vs. Random Overparameterization | The distinction between designed and accidental redundancy |
| Superposition and Polysemanticity | How networks encode more features than neurons by overlapping representations |
| Universal Approximation Theorem | What it proves and doesn't prove about network capability |