← Back to Home← 返回首页

Transformers Are (Naively) Looped Transformers, HorizontallyTransformer 本身就是横向展开的循环 Transformer

Weight sharing across positions makes the Transformer a time-recurrent loop — and TTT is its learned-optimizer cousin跨位置共享权重,让 Transformer 成为一个时间维上的循环结构——而 TTT 是它「把循环体换成梯度下降」的那个近亲

Chunyuan Deng · May 2026Chunyuan Deng · 2026 年 5 月


TL;DR

The standard Transformer is position-invariant: it applies the same weight matrices at every sequence position. This makes it a horizontally looped (time-recurrent) architecture — not just a deep one. The natural generalization is to assign a distinct weight matrix to every position. Test-Time Training (TTT) is exactly this generalization realized through gradient descent: the hidden state stores a fast-weight model that is updated position by position, giving per-position weights without storing a separate matrix for each step.标准 Transformer 是位置不变的:它在每个序列位置上用的是同一组权重矩阵。这让它成为一个横向循环(时间维递归)的架构,而不只是一个「深」的架构。最自然的推广,是给每个位置配一组各自的权重矩阵。Test-Time Training(TTT)正是用梯度下降实现的这种推广:隐状态里存的是一个快权重模型,逐位置更新,于是每个位置都有自己的权重,却不必为每一步单独存一个矩阵。

I. The Standard Transformer, Written Carefully一、把标准 Transformer 仔细写一遍

Fix a sequence of $T$ tokens. After embedding, we have固定一条长度为 $T$ 的 token 序列。做完 embedding 后,我们有

$$X \in \mathbb{R}^{T \times d}$$

where $T$ is the sequence length and $d$ is the model dimension. Denote the $t$-th row as $x_t \in \mathbb{R}^d$.其中 $T$ 是序列长度,$d$ 是模型维度。把第 $t$ 行记作 $x_t \in \mathbb{R}^d$。

A single Transformer layer applies the same projection matrices $W_Q, W_K, W_V, W_O \in \mathbb{R}^{d \times d}$ to every position:单层 Transformer 会把同一组投影矩阵 $W_Q, W_K, W_V, W_O \in \mathbb{R}^{d \times d}$ 作用到每一个位置上:

Self-Attention at position $t$位置 $t$ 上的自注意力
$$q_t = x_t W_Q, \quad k_t = x_t W_K, \quad v_t = x_t W_V \qquad q_t, k_t, v_t \in \mathbb{R}^d$$ $$a_t = \sum_{s \leq t} \frac{\exp(q_t \cdot k_s / \sqrt{d})}{\sum_{s' \leq t} \exp(q_t \cdot k_{s'} / \sqrt{d})}\, v_s \qquad a_t \in \mathbb{R}^d$$ $$o_t = a_t W_O \qquad o_t \in \mathbb{R}^d$$

The crucial observation: $W_Q, W_K, W_V, W_O$ carry no subscript $t$. Every position $t \in \{1, \ldots, T\}$ is processed by the identical linear maps. This is not a coincidence — it is a deliberate design choice called parameter sharing across time.关键的观察是:$W_Q, W_K, W_V, W_O$ 都不带下标 $t$。每个位置 $t \in \{1, \ldots, T\}$ 都被完全相同的线性映射处理。这不是巧合,而是一个刻意的设计选择,叫做沿时间维的参数共享。

The Weight-Sharing Fact权重共享这件事

In a standard Transformer with $L$ layers, the total number of weight parameters is $O(L \cdot d^2)$, independent of sequence length $T$. The same $d^2$ parameters are reused at every one of the $T$ positions within each layer.一个 $L$ 层的标准 Transformer,权重参数总量是 $O(L \cdot d^2)$,与序列长度 $T$ 无关。每层里的这 $d^2$ 个参数,会在全部 $T$ 个位置上被反复使用。

II. Transformers as Horizontally Looped Architectures二、把 Transformer 看作横向循环架构

Think of the computation graph of a single Transformer layer. It has two axes:看一层 Transformer 的计算图,它有两个轴:

Because $W_Q, W_K, W_V, W_O$ are shared across the time axis, the Transformer is a loop unrolled in time:因为 $W_Q, W_K, W_V, W_O$ 在时间轴上是共享的,Transformer 其实是一个在时间上展开的循环:

Transformer as a Horizontal Loop作为横向循环的 Transformer

For each layer $\ell \in \{1,\ldots,L\}$ and position $t \in \{1,\ldots,T\}$:对每一层 $\ell \in \{1,\ldots,L\}$ 和每个位置 $t \in \{1,\ldots,T\}$:

$$h_t^{(\ell)} = f_\theta\!\left(h_t^{(\ell-1)},\; \{h_s^{(\ell-1)}\}_{s \leq t}\right)$$

where $f_\theta$ is the same function (same $\theta = \{W_Q, W_K, W_V, W_O, W_{\text{FF}}\}$) for all $t$.其中对所有 $t$,$f_\theta$ 都是同一个函数(同一组 $\theta = \{W_Q, W_K, W_V, W_O, W_{\text{FF}}\}$)。

This is precisely a recurrent computation along the time axis. The Transformer is not recurrent in the traditional RNN sense (it does not pass a hidden state vector from step to step), but it is parameter-recurrent: the same weight loop is applied at every time step.这恰好就是沿时间轴的递归计算。Transformer 不是传统 RNN 意义上的递归(它不在步与步之间传递隐状态向量),但它是参数意义上的递归:同一组权重构成的循环体,在每个时间步上被重复施加。

Concretely, consider the per-position update at a fixed layer:具体来说,看固定某一层时逐位置的更新:

$$h_t \leftarrow f_\theta(h_t,\; \text{context up to } t), \quad t = 1, 2, \ldots, T$$

This is a loop over $T$ steps, each executing $f_\theta$. The Transformer is therefore a naively looped Transformer — naive in the sense that the loop body $f_\theta$ never changes across iterations.这就是一个跑 $T$ 步的循环,每步执行一次 $f_\theta$。所以 Transformer 是一个朴素的循环 Transformer——朴素之处在于,循环体 $f_\theta$ 在所有迭代中从不改变。

III. The General, Non-Shared Version三、一般形式:不共享权重的版本

The natural generalization removes the weight-sharing constraint: assign a distinct set of parameters to each position.最自然的推广是把权重共享这条约束去掉:给每个位置配一组独立的参数。

Per-Position (Non-Shared) Transformer逐位置(不共享)的 Transformer
$$q_t = x_t W_Q^{(t)}, \quad k_t = x_t W_K^{(t)}, \quad v_t = x_t W_V^{(t)} \qquad W_Q^{(t)}, W_K^{(t)}, W_V^{(t)} \in \mathbb{R}^{d \times d}$$

Here the superscript $(t)$ indicates that each position $t$ has its own projection matrices. The total parameter count is now $O(T \cdot d^2)$ per layer — linear in sequence length.这里的上标 $(t)$ 表示每个位置 $t$ 都有自己的投影矩阵。此时每层的参数量变成 $O(T \cdot d^2)$——随序列长度线性增长。

For $T = 1{,}000{,}000$ and $d = 4{,}096$, this is $\approx 1.7 \times 10^{13}$ parameters per layer — obviously infeasible to store explicitly. Yet this is the conceptually correct, most expressive model: position 1 should arguably use different processing logic than position 1,000,000.取 $T = 1{,}000{,}000$、$d = 4{,}096$,每层约 $1.7 \times 10^{13}$ 个参数——显式存下来显然不现实。但从概念上讲,这才是最正确、表达力最强的模型:位置 1 本来就应该用和位置 1,000,000 不一样的处理逻辑。

The standard Transformer is the special case $W^{(1)} = W^{(2)} = \cdots = W^{(T)} = W$: a single shared matrix used for all positions. Positional encodings are the only mechanism that breaks this symmetry, but they act on the inputs, not the weights.标准 Transformer 是其中的特例 $W^{(1)} = W^{(2)} = \cdots = W^{(T)} = W$:所有位置共用一个矩阵。位置编码是唯一打破这种对称性的机制,但它作用在输入上,而不是权重上。

Standard Transformer标准 Transformer

Weights: $W^{(t)} = W$ for all $t$权重:对所有 $t$ 都有 $W^{(t)} = W$

Parameters: $O(d^2)$ per layer参数量:每层 $O(d^2)$

Position awareness: via positional encodings on inputs位置感知:靠作用在输入上的位置编码

Expressivity: same function applied everywhere表达力:处处都是同一个函数

Per-Position (General) Model逐位置(一般)模型

Weights: $W^{(t)}$ distinct for each $t$权重:每个 $t$ 对应不同的 $W^{(t)}$

Parameters: $O(T \cdot d^2)$ per layer参数量:每层 $O(T \cdot d^2)$

Position awareness: baked into the weights themselves位置感知:直接刻进权重本身

Expressivity: different function at every position表达力:每个位置一个不同的函数

The gap between these two extremes is enormous. Can we find a tractable middle ground that achieves position-dependent processing without storing $T$ full weight matrices?这两个极端之间的差距非常大。有没有一个可行的中间地带,既能做到依位置而变的处理,又不用存下 $T$ 个完整的权重矩阵?

IV. TTT: Horizontal Looping with Learned Gradient Steps四、TTT:把梯度步当作循环体的横向循环

Test-Time Training (TTT) [Sun et al., 2024] closes this gap by replacing the static weight $W$ with a fast-weight model $W_t$ that is updated at each position via gradient descent. The key idea is that instead of storing a per-position weight matrix, we generate it on the fly by running a small optimization.Test-Time Training(TTT)[Sun et al., 2024] 填上了这个空隙:它把静态权重 $W$ 换成一个快权重模型 $W_t$,在每个位置上用梯度下降更新一次。核心想法是,与其把逐位置的权重矩阵存下来,不如靠跑一个小优化现场生成它。

Setup: The Fast-Weight Model设定:快权重模型

At each position $t$, TTT maintains a weight matrix $W_t \in \mathbb{R}^{d \times d}$ (or a small neural network parameterized by $W_t$). The hidden state of the sequence model is $W_t$.在每个位置 $t$,TTT 维护一个权重矩阵 $W_t \in \mathbb{R}^{d \times d}$(或者一个以 $W_t$ 为参数的小网络)。序列模型的隐状态就是 $W_t$。

For each incoming token $x_t \in \mathbb{R}^d$, TTT constructs a self-supervised task:对每个进来的 token $x_t \in \mathbb{R}^d$,TTT 构造一个自监督任务:

TTT Self-Supervised Loss at Position $t$位置 $t$ 上的 TTT 自监督损失
$$\mathcal{L}_t(W) = \left\| W \cdot k_t - v_t \right\|^2$$

where $k_t = x_t W_K \in \mathbb{R}^d$ is the key and $v_t = x_t W_V \in \mathbb{R}^d$ is the target value, with $W_K, W_V \in \mathbb{R}^{d \times d}$ being outer (slow) weights shared across all positions.其中 $k_t = x_t W_K \in \mathbb{R}^d$ 是 key,$v_t = x_t W_V \in \mathbb{R}^d$ 是目标 value,而 $W_K, W_V \in \mathbb{R}^{d \times d}$ 是所有位置共享的外层(慢)权重。

This is a simple ridge regression objective: the fast-weight matrix $W_t$ is asked to map the current key to the current value.这就是一个简单的岭回归目标:要求快权重矩阵 $W_t$ 把当前的 key 映射到当前的 value。

The Recurrent Update Rule递归更新规则

The weight is updated by a gradient step:权重用一步梯度更新:

TTT Update (one gradient step)TTT 更新(一步梯度)
$$W_t = W_{t-1} - \eta \,\nabla_W \mathcal{L}_t(W_{t-1})$$ $$= W_{t-1} - \eta \left(W_{t-1} k_t - v_t\right) k_t^\top$$

where $\eta > 0$ is a learned step size (scalar or per-parameter), and the gradient $\nabla_W \mathcal{L}_t(W) = (Wk_t - v_t)k_t^\top \in \mathbb{R}^{d \times d}$ is the outer product of the residual and the key.其中 $\eta > 0$ 是学出来的步长(标量或逐参数),梯度 $\nabla_W \mathcal{L}_t(W) = (Wk_t - v_t)k_t^\top \in \mathbb{R}^{d \times d}$ 是残差与 key 的外积。

The output at position $t$ is then read out by querying the updated model:位置 $t$ 的输出,就是拿 query 去问更新后的模型:

$$z_t = W_t \cdot q_t \qquad z_t \in \mathbb{R}^d$$

where $q_t = x_t W_Q \in \mathbb{R}^d$ is the query, again using a shared slow weight $W_Q \in \mathbb{R}^{d \times d}$.其中 $q_t = x_t W_Q \in \mathbb{R}^d$ 是 query,同样用共享的慢权重 $W_Q \in \mathbb{R}^{d \times d}$ 得到。

Why This is Per-Position Weights为什么这就是逐位置权重

After $t$ gradient steps, the fast weight $W_t$ encodes the accumulated gradient information from all positions $1, \ldots, t$:走完 $t$ 步梯度之后,快权重 $W_t$ 编码了位置 $1, \ldots, t$ 上累积的全部梯度信息:

$$W_t = W_0 - \eta \sum_{s=1}^{t} \left(W_{s-1} k_s - v_s\right) k_s^\top$$

Each $W_t$ is distinct — it depends on the entire history $(x_1, \ldots, x_t)$. This is precisely the per-position weight matrix $W^{(t)}$ from the general model, but expressed implicitly through gradient accumulation rather than stored explicitly.每个 $W_t$ 都不一样——它依赖于整段历史 $(x_1, \ldots, x_t)$。这正是一般模型里的逐位置权重矩阵 $W^{(t)}$,只不过是用梯度累积隐式表达出来的,而不是显式存下来的。

TTT = Horizontal Looping + Gradient Descent as the Loop BodyTTT = 横向循环 + 以梯度下降为循环体

The standard Transformer loops the same static function $f_W$ over positions. TTT loops a function whose weights change at each step via gradient descent. The loop body is not $f_W$ but $f_{W_t}$, and $W_t$ evolves by a local optimization step at each position.标准 Transformer 在各个位置上循环同一个静态函数 $f_W$。TTT 循环的则是一个权重会变的函数,每步都由梯度下降更新。循环体不是 $f_W$ 而是 $f_{W_t}$,而 $W_t$ 在每个位置上都经过一次局部优化。

This gives TTT the expressivity of per-position weights without the $O(T \cdot d^2)$ storage cost. The trade-off: the per-position weight is determined by a fixed optimization trajectory (gradient descent from $W_0$), not a freely learned mapping.这让 TTT 拿到了逐位置权重的表达力,却不必付出 $O(T \cdot d^2)$ 的存储代价。代价在于:逐位置的权重由一条固定的优化轨迹决定(从 $W_0$ 出发的梯度下降),而不是一个可以自由学习的映射。

V. The LocoProp Connection五、与 LocoProp 的联系

The TTT update is more than vanilla gradient descent — it is closely related to LocoProp [Amid et al., 2022], a local propagation algorithm for training neural networks layer by layer.TTT 的更新不止是普通梯度下降——它和 LocoProp [Amid et al., 2022] 关系密切,后者是一种逐层训练神经网络的局部传播算法。

LocoProp decomposes the global training objective into local per-layer targets. Each layer is trained to minimize:LocoProp 把全局训练目标拆成每层的局部目标。每一层要最小化的是:

LocoProp Local ObjectiveLocoProp 的局部目标
$$\mathcal{L}^{\text{loco}}_\ell(W_\ell) = \left\| W_\ell \cdot h_{\ell-1} - \hat{h}_\ell \right\|^2 + \lambda \left\| W_\ell - W_\ell^{\text{prev}} \right\|^2_F$$

where $h_{\ell-1} \in \mathbb{R}^d$ is the input to layer $\ell$, $\hat{h}_\ell \in \mathbb{R}^d$ is the local target (a detached signal from the next layer's gradient), $W_\ell^{\text{prev}}$ is the previous iterate, and $\|\cdot\|_F$ denotes the Frobenius norm.其中 $h_{\ell-1} \in \mathbb{R}^d$ 是第 $\ell$ 层的输入,$\hat{h}_\ell \in \mathbb{R}^d$ 是局部目标(来自下一层梯度、已 detach 的信号),$W_\ell^{\text{prev}}$ 是上一次迭代的权重,$\|\cdot\|_F$ 表示 Frobenius 范数。

TTT's self-supervised loss $\mathcal{L}_t(W) = \|W k_t - v_t\|^2$ is structurally identical, with the key $k_t$ playing the role of the layer input and the value $v_t$ as the local target. TTT is therefore LocoProp applied horizontally across time rather than vertically across depth.TTT 的自监督损失 $\mathcal{L}_t(W) = \|W k_t - v_t\|^2$ 在结构上与之完全一致:key $k_t$ 扮演层输入的角色,value $v_t$ 充当局部目标。所以 TTT 就是把 LocoProp 横着沿时间轴用了一遍,而不是纵向沿深度用。

LocoProp (Vertical)LocoProp(纵向)

Axis: depth (layer $\ell$)轴:深度(第 $\ell$ 层)

Target: signal from adjacent layer目标:来自相邻层的信号

Update: local regression per layer更新:每层一次局部回归

Purpose: credit assignment without full backprop目的:不做完整反向传播的信用分配

TTT (Horizontal)TTT(横向)

Axis: time (position $t$)轴:时间(位置 $t$)

Target: value $v_t$ at current position目标:当前位置的 value $v_t$

Update: local regression per position更新:每个位置一次局部回归

Purpose: per-position weight adaptation目的:逐位置的权重自适应

VI. Connecting Everything: A Unified View六、串起来:一个统一视角

We can now arrange these architectures on a spectrum of how much weight sharing they enforce along the time axis:现在可以把这些架构按「沿时间轴共享多少权重」排成一个谱:

Spectrum of Time-Axis Weight Sharing时间轴权重共享谱

Full sharing (Standard Transformer): $W^{(t)} = W$ for all $t \in \{1,\ldots,T\}$. One matrix, reused everywhere. Cheapest, but position-invariant.完全共享(标准 Transformer):对所有 $t \in \{1,\ldots,T\}$ 都有 $W^{(t)} = W$。一个矩阵,到处复用。最省,但位置不变。

Gradient-accumulated sharing (TTT): $W^{(t)} = W_0 - \eta\sum_{s \leq t} g_s$, where $g_s$ is the gradient at step $s$. Position-dependent, but constrained to a gradient descent trajectory. $O(d^2)$ hidden state.梯度累积式共享(TTT):$W^{(t)} = W_0 - \eta\sum_{s \leq t} g_s$,其中 $g_s$ 是第 $s$ 步的梯度。依位置而变,但被限制在一条梯度下降轨迹上。隐状态规模 $O(d^2)$。

No sharing (Per-Position Model): $W^{(t)}$ freely chosen for each $t$. Most expressive, but $O(T \cdot d^2)$ parameters — infeasible for long sequences.完全不共享(逐位置模型):每个 $t$ 的 $W^{(t)}$ 自由取值。表达力最强,但参数量 $O(T \cdot d^2)$——长序列下不可行。

The standard Transformer's position-invariance is its defining limitation at long context. When processing position $t = 1{,}000{,}000$, the model uses the exact same weight matrices as at position $t = 1$. The only differentiation comes from the attention pattern over context, not from the processing function itself.位置不变性,是标准 Transformer 在长上下文下最本质的局限。处理位置 $t = 1{,}000{,}000$ 时,模型用的还是位置 $t = 1$ 那组权重矩阵。唯一的区别来自注意力在上下文上的分布,而不是处理函数本身。

TTT breaks this invariance in a principled, memory-efficient way. The fast weight $W_t$ at position $t$ encodes a compressed summary of all keys and values seen so far, updated via a gradient step that costs only $O(d^2)$ per position — the same asymptotic cost as a standard attention operation.TTT 以一种有原则且省显存的方式打破了这种不变性。位置 $t$ 上的快权重 $W_t$ 压缩地记录了此前见过的所有 key 与 value,每个位置只需一步 $O(d^2)$ 的梯度更新——和标准注意力操作是同一个渐近量级。

VII. Summary七、小结

References参考文献

Cite this post引用本文
@misc{deng2026looped,
  author       = {Chunyuan Deng},
  title        = {Transformers Are (Naively) Looped Transformers, Horizontally},
  year         = {2026},
  url          = {https://charlesdddd.github.io/blog/transformers-are-looped.html}
}