Input Gate
Long Short-Term Memory
Easy
Problem
Implement the two LSTM input branches. The input gate chooses how much new information to write, while the candidate branch proposes the new content.
i_t=\sigma\left(W_i[h_{t-1},x_t]+b_i\right).
\widetilde{C}_t=\tanh\left(W_c[h_{t-1},x_t]+b_c\right).
Return a dictionary containing input_gate and candidate_state as float64 NumPy arrays. Each array has shape (H,) for one sample or (N, H) for a batch.
Theory
The input gate and candidate memory form the "write" mechanism of the LSTM cell. While the forget gate decides what to erase from memory, the input gate and candidate together decide what new information to store. They are two separate computations: the input gate produces a vector in [0, 1] controlling how much to write, and the candidate produces a vector in [-1, 1] proposing what to write. Their element-wise product determines the actual update added to the cell state.
What It Is
The input gate and candidate memory are a two-part mechanism that together determine the new information written into the LSTM cell state at each time step. The mechanism is split into two components because "what to write" and "how much to write" require different mathematical properties:
- Input gate i_t: A sigmoid-activated vector in [0, 1]. Each element acts as a valve controlling the magnitude of the update for that cell state dimension. A value near 1 means "write fully," near 0 means "block this update."
- Candidate memory \tilde{C}_t: A tanh-activated vector in [-1, 1]. Each element proposes a new value to add to the cell state. Positive values push the cell state up, negative values push it down.
Both components receive the same concatenated input [h_{t-1}, x_t] but use separate weight matrices and biases. The gate learns to detect when information is relevant, while the candidate learns to encode what that information is. Their element-wise product i_t \odot \tilde{C}_t produces the update vector added to the cell state after the forget gate has done its work.
Key Equations
Step 1 - Concatenate inputs. The previous hidden state h_{t-1} \in \mathbb{R}^{h} and current input x_t \in \mathbb{R}^{d} are concatenated:
[h_{t-1}, x_t] \in \mathbb{R}^{h + d}
This concatenation is shared across all LSTM gates. The combined vector carries both recurrent context and fresh input information.
Step 2 - Input gate. The concatenated vector is transformed by a learned weight matrix and bias, then passed through sigmoid:
i_t = \sigma(W_i \cdot [h_{t-1}, x_t] + b_i)
where W_i \in \mathbb{R}^{h \times (h+d)} and b_i \in \mathbb{R}^{h}. The sigmoid \sigma(z) = \frac{1}{1 + e^{-z}} squashes each element to (0, 1), making the output interpretable as independent "write intensity" controls.
Step 3 - Candidate memory. The same concatenated vector is transformed by a different weight matrix and bias, then passed through tanh:
\tilde{C}_t = \tanh(W_C \cdot [h_{t-1}, x_t] + b_C)
where W_C \in \mathbb{R}^{h \times (h+d)} and b_C \in \mathbb{R}^{h}. The tanh maps each element to (-1, 1), meaning the candidate can propose both increases and decreases to the cell state.
Step 4 - Gated update. The input gate and candidate combine via element-wise multiplication:
i_t \odot \tilde{C}_t \in \mathbb{R}^{h}
Each dimension of the candidate is independently scaled by the corresponding gate value. This product is added to the cell state in the full update equation C_t = f_t \odot C_{t-1} + i_t \odot \tilde{C}_t.
Why Two Separate Components
A natural question is: why not use a single computation to produce the cell state update? The separation exists because magnitude control and content generation are fundamentally different tasks that benefit from different activations and learned representations.
The gate controls magnitude. The input gate i_t operates in [0, 1] and determines the scale of the update per dimension. It answers "how much should I modify this dimension of memory right now?" During training, the gate learns to recognize contexts where memory should be updated (gate near 1) versus contexts where the cell state should be left alone (gate near 0).
The candidate controls content. The candidate \tilde{C}_t operates in [-1, 1] and determines the direction and content of the update. It encodes the actual semantic information from the current input and context into a form suitable for long-term storage.
The element-wise product combines both. The product i_t \odot \tilde{C}_t allows fine-grained control: one dimension might have a strong candidate value of 0.95 but a gate value of only 0.1, meaning the model has encoded useful information but decided now is not the time to write it. Another dimension might have a moderate candidate of 0.4 but a fully open gate of 0.98.
If a single computation produced the update directly (say, a single tanh layer), the model would have no mechanism to selectively suppress individual dimensions. It could only learn to produce small values, which conflates "this information is not relevant now" with "this information has small magnitude." The gate-candidate separation disentangles relevance from content.
Why Tanh for the Candidate
Bounded output in [-1, 1]. The tanh function prevents the candidate from proposing unbounded updates that could cause the cell state to explode. Combined with the gate's [0, 1] range, the maximum possible update to any cell state dimension per time step is bounded between -1 and +1.
Zero-centered output. Unlike sigmoid which outputs in (0, 1) with a mean around 0.5, tanh is centered at zero. This is essential because the candidate needs to both increase and decrease the cell state. If the candidate used sigmoid (always positive), the cell state could only grow over time. Zero-centered outputs let the model naturally push cell state dimensions in either direction.
Stronger gradients near zero. The derivative of tanh at zero is 1 (compared to 0.25 for sigmoid). Since many pre-activation values cluster around zero during training, tanh provides stronger gradient signal for the candidate's weight updates, helping the model learn useful representations faster.
Contrast with sigmoid for gates. Gates use sigmoid because they need to be in [0, 1] to function as multiplicative switches. A gate value of 0 means "block completely" and 1 means "pass completely." Tanh values of -1 and +1 have no such interpretation for a gate. The functions are chosen to match the mathematical role: sigmoid for gating (scaling), tanh for content (representation).
Paper Context: Hochreiter and Schmidhuber (1997)
The input gate was introduced in the original LSTM paper, "Long Short-Term Memory" by Sepp Hochreiter and Jurgen Schmidhuber (1997). The paper describes the input gate's role with a specific framing: "The input gate protects the memory contents from perturbation by irrelevant inputs." This frames the gate as a protective mechanism, not merely a write controller.
The core problem being solved was the vanishing gradient problem in recurrent neural networks. Standard RNNs could not learn long-range dependencies because gradient signals decayed exponentially through time. The LSTM's solution was the constant error carousel (CEC): a cell state updated additively rather than multiplicatively, allowing gradients to flow unchanged across many time steps.
The input gate is essential to making the CEC work. Without it, every input at every time step would perturb the cell state, drowning out stored long-term information. The gate acts as a bouncer: it learns which inputs carry information worth storing and blocks everything else. The paper's use of "protection" emphasizes that the default behavior should be to not write. The cell state is precious, and the input gate must actively decide to open before any modification occurs.
The candidate memory (called "cell input" in the original paper) represents the transformation of current input into a form suitable for storage. The original paper used a squashing function g for this role, which in modern implementations is tanh.
Notably, the original 1997 LSTM did not include a forget gate. The cell state could only be written to (via input gate) and read from (via output gate), but old information could never be explicitly erased. The forget gate was added by Gers et al. (2000). In the original architecture, the input gate was even more critical: once information entered the cell state, it stayed indefinitely, so the gate had to be highly selective because there was no mechanism to undo a bad write.
Numerical Example
Consider an LSTM with hidden size h = 2 and input size d = 2.
Given values:
h_{t-1} = \begin{pmatrix} 0.5 \\ -0.3 \end{pmatrix}, \quad x_t = \begin{pmatrix} 0.8 \\ 0.1 \end{pmatrix}
W_i = \begin{pmatrix} 0.4 & -0.2 & 0.3 & 0.1 \\ 0.1 & 0.5 & -0.1 & 0.6 \end{pmatrix}, \quad b_i = \begin{pmatrix} 0.1 \\ -0.2 \end{pmatrix}
W_C = \begin{pmatrix} -0.3 & 0.7 & 0.2 & -0.4 \\ 0.6 & -0.1 & 0.5 & 0.3 \end{pmatrix}, \quad b_C = \begin{pmatrix} 0.0 \\ 0.1 \end{pmatrix}
Step 1 - Concatenate: [h_{t-1}, x_t] = (0.5, -0.3, 0.8, 0.1)^T
Step 2 - Input gate pre-activation:
Dim 1: (0.4)(0.5) + (-0.2)(-0.3) + (0.3)(0.8) + (0.1)(0.1) + 0.1 = 0.2 + 0.06 + 0.24 + 0.01 + 0.1 = 0.61
Dim 2: (0.1)(0.5) + (0.5)(-0.3) + (-0.1)(0.8) + (0.6)(0.1) + (-0.2) = 0.05 - 0.15 - 0.08 + 0.06 - 0.2 = -0.32
Step 3 - Apply sigmoid:
i_t = \begin{pmatrix} \sigma(0.61) \\ \sigma(-0.32) \end{pmatrix} = \begin{pmatrix} 0.648 \\ 0.421 \end{pmatrix}
Dimension 1 allows about 65% of the candidate through. Dimension 2 allows about 42%.
Step 4 - Candidate pre-activation:
Dim 1: (-0.3)(0.5) + (0.7)(-0.3) + (0.2)(0.8) + (-0.4)(0.1) + 0.0 = -0.15 - 0.21 + 0.16 - 0.04 = -0.24
Dim 2: (0.6)(0.5) + (-0.1)(-0.3) + (0.5)(0.8) + (0.3)(0.1) + 0.1 = 0.3 + 0.03 + 0.4 + 0.03 + 0.1 = 0.86
Step 5 - Apply tanh:
\tilde{C}_t = \begin{pmatrix} \tanh(-0.24) \\ \tanh(0.86) \end{pmatrix} = \begin{pmatrix} -0.235 \\ 0.697 \end{pmatrix}
The candidate proposes a negative update for dimension 1 and a positive update for dimension 2. This illustrates how tanh's zero-centered output naturally allows both directions.
Step 6 - Element-wise product i_t \odot \tilde{C}_t:
i_t \odot \tilde{C}_t = \begin{pmatrix} 0.648 \times (-0.235) \\ 0.421 \times 0.697 \end{pmatrix} = \begin{pmatrix} -0.152 \\ 0.293 \end{pmatrix}
Key observations:
- Dimension 1: The candidate proposed -0.235 but the gate allowed 65% through, resulting in -0.152. The cell state will decrease by this amount.
- Dimension 2: The candidate proposed 0.697 but the gate allowed only 42% through, resulting in $$. Despite a stronger candidate, the lower gate value significantly reduced the update.
- Independence: The gate and candidate are independent. Dimension 2 has a strong candidate but weak gate; dimension 1 has a weaker candidate but stronger gate. Content and relevance are separate judgments.
Connection to GRU
The Gated Recurrent Unit (GRU), introduced by Cho et al. (2014), has its own candidate computation that differs from the LSTM candidate in an important way.
LSTM candidate uses the raw hidden state:
\tilde{C}_t = \tanh(W_C \cdot [h_{t-1}, x_t] + b_C)
GRU candidate uses a reset-gated hidden state:
\tilde{h}_t = \tanh(W \cdot [r_t \odot h_{t-1}, x_t] + b)
where r_t = \sigma(W_r \cdot [h_{t-1}, x_t] + b_r) is the reset gate. Before the hidden state enters the candidate computation, each dimension is scaled by the reset gate. When r_t is near 0, the candidate ignores previous state and computes purely from current input, allowing the GRU to "start fresh."
- LSTM: Relies on the forget gate to manage old information and the input gate to control writing. The candidate always sees the full hidden state.
- GRU: Uses the reset gate to filter the hidden state before it enters the candidate. This makes the candidate conditionally independent of history, something the LSTM candidate cannot do.
The GRU also merges forget and input gates into a single update gate z_t, using the complementary relationship z_t and (1 - z_t): whatever fraction of old state is kept, the complement is used for the new candidate. The LSTM treats forgetting and writing as independent decisions with separate gates, which is more expressive but uses more parameters.
Pitfalls
Using sigmoid instead of tanh for the candidate. If the candidate uses sigmoid, its output is in (0, 1) and can never produce a negative update. The cell state can only increase, destroying the model's ability to push dimensions downward. Tanh's (-1, 1) range is essential for bidirectional updates.
Sharing weight matrices between i_t and \tilde{C}_t. The input gate and candidate must have separate weight matrices (W_i and W_C). Sharing makes gate and candidate deterministic functions of each other, eliminating independent control. The model loses the ability to say "I see important content but now is not the time to write it."
Forgetting these are TWO separate computations returning a tuple. A common error is treating the input gate as a single operation. You must compute both i_t and \tilde{C}_t as separate tensors, then combine them with element-wise multiplication. If your implementation fuses them into a single linear layer followed by one activation, the mechanism is broken.
Incorrect concatenation order. The convention [h_{t-1}, x_t] means hidden state comes first, then input. If W_i is shaped h \times (h + d), the first h columns multiply h_{t-1} and the last d columns multiply x_t. Reversing concatenation without transposing the weight matrix produces wrong results.
Confusing \tilde{C}_t with C_t. The candidate \tilde{C}_t (with tilde) is the proposed update. The actual cell state C_t (without tilde) is the result of C_t = f_t \odot C_{t-1} + i_t \odot \tilde{C}_t. These are different tensors at different stages. The input gate problem only asks you to compute i_t and \tilde{C}_t, not the final C_t.
Applying the gate before the activation. The correct order is: linear transform, then activation, for each component separately. If you apply the gate to pre-activation values or compute \sigma(\tanh(W \cdot [h, x])), the math is completely different and the mechanism breaks.
Examples
Example 1
- Input
h_prev = [0,0], x_t = [1,0.5], W_i = [[0.2,0.3,0.1,0.4],[0.4,0.2,0.3,0.1]], b_i = [0,0], W_c = [[0.3,0.1,0.2,0.1],[0.1,0.4,0.1,0.3]], b_c = [0,0]- Output
{"input_gate":[0.574443,0.586618],"candidate_state":[0.244919,0.244919]}- Explanation
- Both branches read the same joined features but use different weights and activations.
Example 2
- Input
h_prev = [0.3,-0.2], x_t = [0.5,1], W_i = [[0.2,0.3,0.1,0.4],[0.4,0.2,0.3,0.1]], b_i = [0.1,-0.1], W_c = [[0.3,0.1,0.2,0.1],[0.1,0.4,0.1,0.3]], b_c = [0.05,0.1]- Output
{"input_gate":[0.634136,0.557248],"candidate_state":[0.309507,0.379949]}
Example 3
- Input
h_prev = [[0,0],[0.3,-0.2]], x_t = [[1,0.5],[0.5,1]], W_i = [[0.2,0.3,0.1,0.4],[0.4,0.2,0.3,0.1]], b_i = [0,0], W_c = [[0.3,0.1,0.2,0.1],[0.1,0.4,0.1,0.3]], b_c = [0,0]- Output
{"input_gate":[[0.574443,0.586618],[0.610639,0.581759]],"candidate_state":[[0.244919,0.244919],[0.263625,0.291313]]}
Hints
- Compute joined once and reuse it for both affine transforms.
- Use sigmoid for input_gate and np.tanh for candidate_state.
Requirements
- Use the same [h_prev, x_t] concatenation for both branches.
- Apply sigmoid to the input-gate branch and tanh to the candidate branch.
- Return exactly the keys input_gate and candidate_state.
Constraints
- Unbatched h_prev and x_t have shapes (H,) and (D,); batched inputs have shapes (N, H) and (N, D).
- W_i and W_c have shape (H, H + D).
- b_i and b_c have shape (H,).
- Both returned arrays use float64.
Starter Code
import numpy as np
def sigmoid(x: np.ndarray) -> np.ndarray:
return 1.0 / (1.0 + np.exp(-np.clip(x, -500, 500)))
def input_gate(h_prev: np.ndarray, x_t: np.ndarray,
W_i: np.ndarray, b_i: np.ndarray,
W_c: np.ndarray, b_c: np.ndarray) -> dict:
"""
Returns input_gate and candidate_state as float64 arrays.
"""
passTest Cases
| Case | Matches | |
|---|---|---|
| Basic input gate | — | public |
| input gate with bias | — | public |
| Batched input gate | — | public |