MediumDenseNet

Composite Layer (BN-ReLU-Conv)

DenseNet

Medium

Problem

Implement the composite transformation used by a DenseNet layer. In evaluation mode, apply per-channel batch normalization, ReLU, and a bias-free 3 by 3 convolution in that order.

\widehat{x}_{n,c,h,w}=\gamma_c\frac{x_{n,c,h,w}-\mu_c}{\sqrt{\sigma_c^2+\varepsilon}}+\beta_c

y=\operatorname{Conv}_{3\times3}\!\left(\operatorname{ReLU}(\widehat{x})\right)

Here, x has shape (N,C,H,W), \gamma and \beta are the batch-normalization scale and shift, \mu and \sigma^2 are the supplied running mean and variance, and \varepsilon is eps. The convolution uses stride 1, padding 1, and no bias. Return y as a float64 PyTorch tensor with shape (N,K,H,W), where K is the first dimension of conv_weight.

Theory

The composite layer is the elementary computation unit inside a DenseNet (Huang et al., 2017). It is the function H_\ell(\cdot) that every layer in a dense block applies to its input: a Batch Normalization, then a ReLU nonlinearity, then a 3 \times 3 convolution that produces a small fixed number of new feature maps called the growth rate.


What It Is

In a Densely Connected Convolutional Network, layer \ell receives the concatenation of the feature maps of all preceding layers as its input. It then computes a nonlinear transformation H_\ell and outputs k new feature maps, where k is the growth rate. The paper defines H_\ell as a composite function of three consecutive operations.

Concretely, H_\ell is the ordered composition:

This problem implements the basic (non bottleneck) composite layer exactly as it appears in Section 3 of the paper.


Key Equations

For an input tensor x \in \mathbb{R}^{N \times C \times H \times W}, evaluation mode Batch Normalization standardizes each channel c using the stored running mean \mu_c and running variance \sigma^2_c, then rescales with learned parameters \gamma_c and \beta_c:

\hat{x}_{n,c,i,j} = \frac{x_{n,c,i,j} - \mu_c}{\sqrt{\sigma^2_c + \epsilon}}, \qquad y_{n,c,i,j} = \gamma_c \, \hat{x}_{n,c,i,j} + \beta_c

where \epsilon is a small constant (default 10^{-5}) added inside the square root for numerical stability. The activation is then rectified:

z_{n,c,i,j} = \max(0, \, y_{n,c,i,j})

Finally a 3 \times 3 convolution with weight W \in \mathbb{R}^{k \times C \times 3 \times 3} and no bias produces k output channels with the spatial size preserved by padding = 1:

H_\ell(x)_{n,o,i,j} = \sum_{c=1}^{C} \sum_{a=0}^{2} \sum_{b=0}^{2} W_{o,c,a,b} \; z_{n,c,\, i+a-1,\, j+b-1}

The output shape is (N, k, H, W). The channel count collapses from C to the growth rate k, while the spatial resolution H \times W is unchanged.


The Pre-Activation Order and Why It Matters

The ordering BN \to ReLU \to Conv is called pre-activation, because the normalization and activation come before the convolution rather than after it. DenseNet adopts this order directly from the pre-activation ResNet study (He et al., 2016), which showed that placing BN and ReLU ahead of the weight layer improves gradient flow and gives cleaner identity paths in very deep networks.

Several properties make pre-activation a natural fit for DenseNet:

The original (post-activation) ResNet used Conv \to BN \to ReLU. Swapping to BN \to ReLU \to Conv is not cosmetic: it changes which tensor is normalized and which tensor is the layer output, and it produces numerically different results. Getting this order wrong is the single most common implementation error for this problem.


Growth Rate

The growth rate k is the number of feature maps each composite layer contributes. It is the output channel count of the 3 \times 3 convolution. DenseNet uses a surprisingly small growth rate (the paper reports strong ImageNet results with k = 32, and CIFAR experiments with k = 12), which is one reason the architecture is so parameter efficient.

Because layer \ell receives the concatenation of all earlier outputs, its input channel count is C_0 + (\ell - 1) k, where C_0 is the number of channels entering the block. The input width therefore grows linearly with depth, but each layer only adds k new maps. The small per layer contribution is what keeps the total model compact despite the dense connectivity.

In this problem, the input channel count C is whatever the concatenated input happens to be, and the convolution weight has shape (k, C, 3, 3), so the growth rate is read directly from the first dimension of the weight tensor.


Same Padding and the 3x3 Convolution

A 3 \times 3 kernel with stride 1 and padding 1 is a same convolution: the output spatial dimensions equal the input spatial dimensions. The general formula for one spatial axis is:

H_{\text{out}} = \left\lfloor \frac{H_{\text{in}} + 2p - k_{\text{size}}}{s} \right\rfloor + 1

With k_{\text{size}} = 3, p = 1, s = 1 this gives H_{\text{out}} = H_{\text{in}} + 2 - 3 + 1 = H_{\text{in}}. Preserving spatial size is mandatory inside a dense block: every layer's output must match the spatial dimensions of every other layer's output so that they can be concatenated along the channel axis. If the convolution dropped the padding, the output would shrink by two pixels in each spatial dimension and concatenation would fail.

The convolution carries no bias in this composite layer. The affine shift \beta from the preceding BN already provides a per channel additive degree of freedom, so a separate convolution bias would be redundant. Adding a spurious bias term changes the output and is a common bug.


Evaluation Mode Batch Normalization

During training, BN computes the mean and variance from the current mini batch and updates a running estimate. At inference time the layer is frozen: it uses the stored running statistics instead of batch statistics, so the transformation becomes a fixed, deterministic affine map per channel. This problem uses the eval mode behavior, which is why the running mean \mu_c and running variance \sigma^2_c are supplied as inputs rather than computed from x.

Eval mode is the correct choice for a from scratch composite layer test, because it removes the dependence on batch size and makes the output reproducible. Note the normalization is applied per channel: \mu, \sigma^2, \gamma, and \beta are all vectors of length C, and each is broadcast across the batch and spatial dimensions.

A useful sanity check is the identity BN setting: with \gamma = 1, \beta = 0, \mu = 0, \sigma^2 = 1, the normalization reduces to \hat{x} = x / \sqrt{1 + \epsilon}, which is approximately x for small \epsilon. In this case the composite layer is essentially ReLU followed by the convolution.


Paper Context and Design Decisions

The paper motivates the composite layer with a single goal: maximize information and gradient flow between layers. Where ResNet adds the outputs of layers through summation, DenseNet concatenates them. Summation can impede the flow of information because the identity and the residual are combined into one tensor, whereas concatenation keeps every feature map distinct and accessible. The composite function H_\ell is the operator that turns the accumulated, concatenated state into a fresh set of features.

Three explicit design decisions in Section 3 shape this layer:

These choices together explain why the composite layer looks the way it does: a normalization that can cope with heterogeneous concatenated inputs, an activation, and a small same padded convolution that emits exactly k new maps.


Pre-Activation vs Post-Activation

It is worth contrasting the two orderings directly, since they are easy to confuse:

The difference matters most in deep networks with many shortcut connections. In the pre-activation form, the path that a feature map takes to a downstream layer is not interrupted by a normalization or activation that depends on the addition or concatenation, so gradients propagate more cleanly. For DenseNet, where each feature map may be read by many later layers, keeping the output un-activated also means each consumer can normalize the channels in the way that suits its own computation.

A practical consequence: in the pre-activation form the very first layer of a block normalizes its raw input, and the very last operation of every layer is a linear convolution. If an implementation accidentally appends a trailing ReLU after the convolution, the outputs become non negative and concatenation feeds only rectified features forward, which is not what the paper specifies.


Role Inside a Dense Block

A dense block stacks several composite layers. Layer \ell takes the concatenation [x_0, x_1, \ldots, x_{\ell-1}] of all previous feature maps, applies H_\ell, and produces x_\ell, a tensor of exactly k channels. That output is then concatenated onto the running collection for the next layer to consume.

This dense connectivity is the defining idea of the paper: every layer has direct access to the feature maps of all layers before it, and its own output is passed to all layers after it. The composite layer is the per step computation that makes this possible. It must therefore be self contained, normalize its (possibly heterogeneous) input, and emit a small, spatially aligned set of new features.

The deeper bottleneck variant (DenseNet-B) inserts a 1 \times 1 convolution before the 3 \times 3 convolution to reduce input channels first. This problem implements the basic composite layer without the bottleneck, matching the simplest H_\ell definition in the paper.


Worked Example (N=1, C=2, H=W=2, k=1)

Take a single channel pair with x[0,0] = \begin{pmatrix} 1 & -1 \ 0 & 2 \end{pmatrix} and x[0,1] = \begin{pmatrix} 0 & 1 \ -2 & 1 \end{pmatrix}. Let \gamma = [1, 1], \beta = [0, 0], \mu = [0, 0], \sigma^2 = [1, 1], and \epsilon = 0.

  1. BN (identity here): with these statistics \hat{x} = x and y = x, so the tensors pass through unchanged.

  2. ReLU: clamp negatives to zero. Channel 0 becomes \begin{pmatrix} 1 & 0 \ 0 & 2 \end{pmatrix} and channel 1 becomes \begin{pmatrix} 0 & 1 \ 0 & 1 \end{pmatrix}.

  3. Convolution: a 3 \times 3 same padded kernel slides over the zero padded 2 \times 2 activations, summing the element wise products across both input channels at every position. The result is a single 2 \times 2 output map (because k = 1), with spatial size preserved by the padding.

The essential observations: the negatives are removed before the convolution sees them, and the two input channels are fused into one output channel by summing over the channel dimension inside the convolution.


Pitfalls


Examples

Example 1

Input
x = [[[[1,-1],[0.5,2]]]], bn_gamma = [1.2], bn_beta = [-0.1], bn_mean = [0.25], bn_var = [0.5], conv_weight = [[[[0.1,0,-0.1],[0.2,0.5,0.2],[-0.1,0,0.1]]]], eps = 0.00001
Output
[[[[0.873372,0.20213],[0.736094,1.617039]]]]
Explanation
Batch normalization and ReLU prepare the feature map before the supplied same-padding convolution creates the new channels.

Example 2

Input
x.shape = (1, 2, 3, 3), bn_gamma.shape = (2), bn_beta.shape = (2), bn_mean.shape = (2), bn_var.shape = (2), conv_weight.shape = (1, 2, 3, 3), eps = 0.00001
Output
[[[[0.134607,-0.331748,-0.244685],[-0.351103,-0.70827,-0.056557],[-0.033623,-0.116317,-0.159933]]]]

Example 3

Input
x.shape = (2, 1, 2, 3), bn_gamma.shape = (1), bn_beta.shape = (1), bn_mean.shape = (1), bn_var.shape = (1), conv_weight.shape = (2, 1, 3, 3), eps = 0.00001
Output
[[[[0.268063,0.316078,0.130376],[0.117408,-0.105753,-0.196717]],[[-0.312351,-0.189428,0.108743],[0.146834,0.363691,-0.144186]]],[[[0.521671,0.228662,0.051559],[-0.285679,0.140686,0.140373]],[[-0.287335,0.244076,0.09081],[0.086422,-0.051278,-0.044722]]]]

Hints

  1. Reshape each channel vector to (1, C, 1, 1) before normalization.
  2. Use F.conv2d with padding=1 and bias=None.

Requirements

Constraints

Starter Code

import torch
import torch.nn.functional as F

def composite_layer(x: torch.Tensor, bn_gamma: torch.Tensor, bn_beta: torch.Tensor,
                    bn_mean: torch.Tensor, bn_var: torch.Tensor,
                    conv_weight: torch.Tensor, eps: float = 1e-5) -> torch.Tensor:
    """
    Returns the float64 output of the DenseNet composite transformation.
    """
    pass

Test Cases

CaseMatches
Single channelpublic
Two channelspublic
Two outputspublic