EasyAlexNet

AlexNet Convolution Layer

ImageNet Classification with Deep Convolutional Neural Networks

Easy

Problem

Implement the cross-correlation used by an AlexNet-style convolution layer for an NHWC image batch. For image X, kernel K, bias b, stride s, and zero-padding p, each output value is

Y_{n,i,j,o}=b_o+\sum_{u=0}^{K_h-1}\sum_{v=0}^{K_w-1}\sum_{c=0}^{C-1}X^{(p)}_{n,is+u,js+v,c}K_{u,v,c,o}.

Here, n selects a batch item, (i,j) selects an output location, c is an input channel, o is an output channel, and X^{(p)} is the zero-padded image. Return the convolution result as a float64 NumPy array with shape (B,H_{out},W_{out},C_{out}).

Theory

A convolutional layer applies learnable filters to an input tensor, producing feature maps that detect local spatial patterns (edges, textures, shapes). Krizhevsky et al. (2012) used five convolutional layers as the backbone of AlexNet, the model that won ImageNet LSVRC-2012 and reignited modern deep learning.


What a Convolution Layer Does

A convolutional layer slides a small filter (kernel) across an input tensor. At every position, the filter computes a dot product with the input patch it covers, producing a single scalar. Repeating this across all valid positions generates a 2D feature map (or activation map).


The Output Dimension Formula

For each spatial axis:

H_{out} = \lfloor \frac{H_{in} + 2p - k}{s} \rfloor + 1

After padding, effective input size is H_{in} + 2p. The kernel's furthest valid start position is H_{in} + 2p - k. Dividing by stride gives the number of steps; adding 1 accounts for the initial position. The floor discards any leftover pixels that don't fit a complete kernel window.

Full output tensor shape:

(B, H_{out}, W_{out}, F)

where B is the batch size. The width formula is identical with W_{in} substituted for H_{in}.


What Each Parameter Controls

Kernel Size

Kernel size k defines the local neighborhood each filter sees. Larger kernels have a larger receptive field but more parameters (k^2 \times C_{in} per filter). AlexNet uses 11 \times 11 in Conv1 (to capture large-scale edges and color gradients from raw pixels) and progressively smaller kernels ($ \times 5$, then 3 \times 3) in deeper layers where each position already represents a large image region.

Stride

Stride s controls how far the kernel moves between applications. Stride 1 produces maximum output size; stride 4 reduces output size by roughly 4x per spatial dimension. Unlike pooling, strided convolution downsamples while learning features. AlexNet's stride 4 in Conv1 aggressively reduces $ \times 224$ to 55 \times 55 in one layer -- a pragmatic choice for the GPUs available in 2012.

Padding

Padding adds zero-rows and zero-columns around the input border. Without it, each layer shrinks spatial dimensions by k - 1 pixels (at stride 1).

AlexNet uses specific per-layer values: padding 2 for Conv1 and Conv2, padding 1 for Conv3-5.

Number of Filters

The number of filters F determines feature diversity. More filters = more capacity, more parameters. AlexNet uses 96 -> 256 -> 384 -> 384 -> 256 filters across Conv1-5. This pattern of increasing channels while decreasing spatial size is a hallmark of CNN architectures: as resolution drops, the network compensates with more feature channels.


AlexNet's Five Convolutional Layers

AlexNet has five convolutional layers followed by three fully connected layers. The complete convolutional stack:

After Conv5 pooling, the output is flattened to 6 \times 6 \times 256 = 9{,}216 elements feeding the fully connected layers. The conv layers contain most of the computation but a minority of the parameters; the FC layers contain the majority of parameters.


The Paper's Design Decisions

Why 11 \times 11 Kernels in the First Layer

Conv1 operates on raw RGB pixels. An 11 \times 11 kernel covers 121 pixels per channel -- large enough to detect oriented edges, color blobs, and texture gradients. Smaller kernels would have too narrow a view at this stage. Combined with stride 4, adjacent output positions have overlapping receptive fields ($ - 4 = 7$ pixel overlap), downsampling from 224 \times 224 to 55 \times 55 in one layer.

Why Smaller Kernels in Deeper Layers

By Conv3-5, each spatial position's effective receptive field already covers a large image region (each pixel summarizes many original pixels from prior layers). A 3 \times 3 kernel at this depth is sufficient to combine high-level features.

GPU Memory Constraints in 2012

AlexNet was split across two GTX 580 GPUs (3 GB each). Each GPU handled half the filters in most layers. Conv3 is an exception: it receives all 256 channels from both GPUs, while Conv4 and Conv5 connect only within the same GPU. This hardware-specific optimization influenced the filter counts (96, 256, 384, 384, 256) but is no longer relevant with modern GPUs.

The Role of LRN and Pooling Between Layers

Local Response Normalization (LRN) after Conv1 and Conv2 normalizes neuron responses by dividing by summed squared responses of neighbors across adjacent feature maps. The paper reports LRN reduces top-1 error by 1.4% and top-5 by 1.2%. Later architectures (VGGNet, ResNet) found LRN unnecessary and replaced it with Batch Normalization.

Max pooling (3 \times 3, stride 2) after Conv1, Conv2, and Conv5 uses overlapping pooling (pool 3, stride 2). Krizhevsky et al. report this reduces top-1 error by 0.4% and top-5 by 0.3% vs. non-overlapping pooling.


Parameter Count

Parameters per convolutional layer:

\text{params} = F \times (k^2 \times C_{in} + 1)

The +1 accounts for one bias per filter. For AlexNet:

Total conv parameters: ~3.7 million. By comparison, the first FC layer alone has 9{,}216 \times 4{,}096 = 37{,}748{,}736 parameters (roughly 10x all conv layers combined). This illustrates that conv layers are parameter-efficient due to weight sharing across spatial positions.


Computational Cost (FLOPs)

FLOPs per convolutional layer:

\text{FLOPs} \approx 2 \times H_{out} \times W_{out} \times F \times k^2 \times C_{in}

The factor of 2 accounts for multiply-accumulate operations. For Conv1:

\text{FLOPs}_{\text{Conv1}} \approx 2 \times 55 \times 55 \times 96 \times 11^2 \times 3 \approx 211 \text{ million}

Conv layers dominate computation (~90% of total FLOPs) while FC layers dominate parameter count. This asymmetry is characteristic of CNNs.


Numerical Example: Conv1 Shape Computation

Input: batch of 2 RGB images, each 224 \times 224. Tensor shape: (2, 224, 224, 3). Conv1: $ = 11$, s = 4, p = 2, F = 96.

  1. After padding: 224 + 2 \times 2 = 228 per axis.

  2. Valid kernel placements: Span available: 228 - 11 = 217. Steps: \lfloor 217 / 4 \rfloor = 54. Plus starting position: 54 + 1 = 55.

  3. Formula verification: H_{out} = \lfloor \frac{224 + 4 - 11}{4} \rfloor + 1 = \lfloor \frac{217}{4} \rfloor + 1 = 54 + 1 = 55.

  4. Width: Same (square input). W_{out} = 55.

  5. Output channels: C_{out} = 96.

  6. Final output shape: (2, 55, 55, 96).

  7. Parameters: 96 \times (11 \times 11 \times 3 + 1) = 96 \times 364 = 34{,}944.


Numerical Example: Tracing Through All Five Layers

Starting from (1, 224, 224, 3):

Conv1

k = 11, s = 4, p = 2, F = 96.

H_{out} = \lfloor \frac{224 + 4 - 11}{4} \rfloor + 1 = 55

Output: (1, 55, 55, 96). After ReLU, LRN, max pooling (k_p = 3, s_p = 2):

H_{pool} = \lfloor \frac{55 - 3}{2} \rfloor + 1 = 27

After pooling: (1, 27, 27, 96).

Conv2

k = 5, s = 1, p = 2, F = 256.

H_{out} = \lfloor \frac{27 + 4 - 5}{1} \rfloor + 1 = 27

Output: (1, 27, 27, 256). After ReLU, LRN, max pooling (k_p = 3, s_p = 2):

H_{pool} = \lfloor \frac{27 - 3}{2} \rfloor + 1 = 13

After pooling: (1, 13, 13, 256).

Conv3

k = 3, s = 1, p = 1, F = 384.

H_{out} = \lfloor \frac{13 + 2 - 3}{1} \rfloor + 1 = 13

Output: (1, 13, 13, 384). ReLU only.

Conv4

k = 3, s = 1, p = 1, F = 384.

H_{out} = \lfloor \frac{13 + 2 - 3}{1} \rfloor + 1 = 13

Output: (1, 13, 13, 384). ReLU only.

Conv5

k = 3, s = 1, p = 1, F = 256.

H_{out} = \lfloor \frac{13 + 2 - 3}{1} \rfloor + 1 = 13

Output: (1, 13, 13, 256). After ReLU, max pooling (k_p = 3, s_p = 2):

H_{pool} = \lfloor \frac{13 - 3}{2} \rfloor + 1 = 6

After pooling: (1, 6, 6, 256). Flattened to (1, 9216) before FC layers.

Spatial progression: 224 \to 55 \to 27 \to 27 \to 13 \to 13 \to 13 \to 13 \to 6. Aggressive downsampling happens in Conv1 (stride 4) and the three pooling stages. Conv3 and Conv4 use same-padding to preserve spatial size.


From AlexNet to Modern Convolutions

VGGNet: Only 3 \times 3 Kernels

Simonyan and Zisserman (2014) showed that stacking 3 \times 3 layers matches large-kernel receptive fields with fewer parameters and more non-linearity. Two 3 \times 3 layers = $ \times 5$ receptive field with 2 \times 9 = 18 weights vs. 25. Three 3 \times 3 layers = $ \times 11$ field with 27 weights vs. 121.

ResNet: Residual Connections

He et al. (2015) introduced skip connections (\text{output} = F(x) + x), enabling training of 50-152 layer networks by addressing vanishing gradients. ResNet uses $ \times 1$ convolutions for channel adjustment and a 7 \times 7 first-layer kernel at stride 2 followed by $ \times 3$ max pool.

Depthwise Separable Convolutions

MobileNet (Howard et al., 2017) decomposed standard convolution into depthwise (k \times k \times 1 per channel) and pointwise (1 \times 1 \times C_{in}) convolutions. For AlexNet's Conv3: standard = $ \times 3^2 \times 256 = 884{,}736$ MACs per output position; depthwise separable = $ \times 256 + 256 \times 384 = 100{,}608$ (8.8x reduction).


Pitfalls


Examples

Example 1

Input
image = [[[[1],[2],[3]],[[4],[5],[6]],[[7],[8],[9]]]], kernel = [[[[1]],[[0]]],[[[0]],[[-1]]]], bias = [0], stride = 1, padding = 0
Output
[[[[-4],[-4]],[[-4],[-4]]]]
Explanation
Each kernel window is multiplied elementwise with the image patch, summed, and shifted by the output-channel bias.

Example 2

Input
image = [[[[1,2],[3,4]],[[5,6],[7,8]]]], kernel = [[[[2],[-1]]]], bias = [0.5], stride = 1, padding = 0
Output
[[[[0.5],[2.5]],[[4.5],[6.5]]]]

Example 3

Input
image = [[[[1],[2]],[[3],[4]]]], kernel = [[[[1]],[[1]]],[[[1]],[[1]]]], bias = [0], stride = 2, padding = 1
Output
[[[[1],[2]],[[3],[4]]]]

Hints

  1. Use np.pad only on the height and width axes.
  2. np.tensordot can contract a patch over its height, width, and input-channel axes.
  3. The output sizes are floor((size + 2 * padding - kernel_size) / stride) + 1.

Requirements

Constraints

Starter Code

import numpy as np

def alexnet_conv1(image: np.ndarray, kernel: np.ndarray, bias: np.ndarray,
                   stride: int, padding: int) -> np.ndarray:
    """
    Returns the float64 NHWC convolution output.
    """
    pass

Test Cases

CaseMatches
Single-channel valid convolutionpublic
One-by-one two-channel convolutionpublic
Padded strided convolutionpublic