Weight sharing across positions makes the Transformer a time-recurrent loop — and TTT is its learned-optimizer cousin跨位置共享权重,让 Transformer 成为一个时间维上的循环结构——而 TTT 是它「把循环体换成梯度下降」的那个近亲
Fix a sequence of $T$ tokens. After embedding, we have固定一条长度为 $T$ 的 token 序列。做完 embedding 后,我们有
$$X \in \mathbb{R}^{T \times d}$$where $T$ is the sequence length and $d$ is the model dimension. Denote the $t$-th row as $x_t \in \mathbb{R}^d$.其中 $T$ 是序列长度,$d$ 是模型维度。把第 $t$ 行记作 $x_t \in \mathbb{R}^d$。
A single Transformer layer applies the same projection matrices $W_Q, W_K, W_V, W_O \in \mathbb{R}^{d \times d}$ to every position:单层 Transformer 会把同一组投影矩阵 $W_Q, W_K, W_V, W_O \in \mathbb{R}^{d \times d}$ 作用到每一个位置上:
The crucial observation: $W_Q, W_K, W_V, W_O$ carry no subscript $t$. Every position $t \in \{1, \ldots, T\}$ is processed by the identical linear maps. This is not a coincidence — it is a deliberate design choice called parameter sharing across time.关键的观察是:$W_Q, W_K, W_V, W_O$ 都不带下标 $t$。每个位置 $t \in \{1, \ldots, T\}$ 都被完全相同的线性映射处理。这不是巧合,而是一个刻意的设计选择,叫做沿时间维的参数共享。
In a standard Transformer with $L$ layers, the total number of weight parameters is $O(L \cdot d^2)$, independent of sequence length $T$. The same $d^2$ parameters are reused at every one of the $T$ positions within each layer.一个 $L$ 层的标准 Transformer,权重参数总量是 $O(L \cdot d^2)$,与序列长度 $T$ 无关。每层里的这 $d^2$ 个参数,会在全部 $T$ 个位置上被反复使用。
Think of the computation graph of a single Transformer layer. It has two axes:看一层 Transformer 的计算图,它有两个轴:
Because $W_Q, W_K, W_V, W_O$ are shared across the time axis, the Transformer is a loop unrolled in time:因为 $W_Q, W_K, W_V, W_O$ 在时间轴上是共享的,Transformer 其实是一个在时间上展开的循环:
For each layer $\ell \in \{1,\ldots,L\}$ and position $t \in \{1,\ldots,T\}$:对每一层 $\ell \in \{1,\ldots,L\}$ 和每个位置 $t \in \{1,\ldots,T\}$:
$$h_t^{(\ell)} = f_\theta\!\left(h_t^{(\ell-1)},\; \{h_s^{(\ell-1)}\}_{s \leq t}\right)$$where $f_\theta$ is the same function (same $\theta = \{W_Q, W_K, W_V, W_O, W_{\text{FF}}\}$) for all $t$.其中对所有 $t$,$f_\theta$ 都是同一个函数(同一组 $\theta = \{W_Q, W_K, W_V, W_O, W_{\text{FF}}\}$)。
This is precisely a recurrent computation along the time axis. The Transformer is not recurrent in the traditional RNN sense (it does not pass a hidden state vector from step to step), but it is parameter-recurrent: the same weight loop is applied at every time step.这恰好就是沿时间轴的递归计算。Transformer 不是传统 RNN 意义上的递归(它不在步与步之间传递隐状态向量),但它是参数意义上的递归:同一组权重构成的循环体,在每个时间步上被重复施加。
Concretely, consider the per-position update at a fixed layer:具体来说,看固定某一层时逐位置的更新:
$$h_t \leftarrow f_\theta(h_t,\; \text{context up to } t), \quad t = 1, 2, \ldots, T$$This is a loop over $T$ steps, each executing $f_\theta$. The Transformer is therefore a naively looped Transformer — naive in the sense that the loop body $f_\theta$ never changes across iterations.这就是一个跑 $T$ 步的循环,每步执行一次 $f_\theta$。所以 Transformer 是一个朴素的循环 Transformer——朴素之处在于,循环体 $f_\theta$ 在所有迭代中从不改变。
The natural generalization removes the weight-sharing constraint: assign a distinct set of parameters to each position.最自然的推广是把权重共享这条约束去掉:给每个位置配一组独立的参数。
Here the superscript $(t)$ indicates that each position $t$ has its own projection matrices. The total parameter count is now $O(T \cdot d^2)$ per layer — linear in sequence length.这里的上标 $(t)$ 表示每个位置 $t$ 都有自己的投影矩阵。此时每层的参数量变成 $O(T \cdot d^2)$——随序列长度线性增长。
For $T = 1{,}000{,}000$ and $d = 4{,}096$, this is $\approx 1.7 \times 10^{13}$ parameters per layer — obviously infeasible to store explicitly. Yet this is the conceptually correct, most expressive model: position 1 should arguably use different processing logic than position 1,000,000.取 $T = 1{,}000{,}000$、$d = 4{,}096$,每层约 $1.7 \times 10^{13}$ 个参数——显式存下来显然不现实。但从概念上讲,这才是最正确、表达力最强的模型:位置 1 本来就应该用和位置 1,000,000 不一样的处理逻辑。
The standard Transformer is the special case $W^{(1)} = W^{(2)} = \cdots = W^{(T)} = W$: a single shared matrix used for all positions. Positional encodings are the only mechanism that breaks this symmetry, but they act on the inputs, not the weights.标准 Transformer 是其中的特例 $W^{(1)} = W^{(2)} = \cdots = W^{(T)} = W$:所有位置共用一个矩阵。位置编码是唯一打破这种对称性的机制,但它作用在输入上,而不是权重上。
Weights: $W^{(t)} = W$ for all $t$权重:对所有 $t$ 都有 $W^{(t)} = W$
Parameters: $O(d^2)$ per layer参数量:每层 $O(d^2)$
Position awareness: via positional encodings on inputs位置感知:靠作用在输入上的位置编码
Expressivity: same function applied everywhere表达力:处处都是同一个函数
Weights: $W^{(t)}$ distinct for each $t$权重:每个 $t$ 对应不同的 $W^{(t)}$
Parameters: $O(T \cdot d^2)$ per layer参数量:每层 $O(T \cdot d^2)$
Position awareness: baked into the weights themselves位置感知:直接刻进权重本身
Expressivity: different function at every position表达力:每个位置一个不同的函数
The gap between these two extremes is enormous. Can we find a tractable middle ground that achieves position-dependent processing without storing $T$ full weight matrices?这两个极端之间的差距非常大。有没有一个可行的中间地带,既能做到依位置而变的处理,又不用存下 $T$ 个完整的权重矩阵?
Test-Time Training (TTT) [Sun et al., 2024] closes this gap by replacing the static weight $W$ with a fast-weight model $W_t$ that is updated at each position via gradient descent. The key idea is that instead of storing a per-position weight matrix, we generate it on the fly by running a small optimization.Test-Time Training(TTT)[Sun et al., 2024] 填上了这个空隙:它把静态权重 $W$ 换成一个快权重模型 $W_t$,在每个位置上用梯度下降更新一次。核心想法是,与其把逐位置的权重矩阵存下来,不如靠跑一个小优化现场生成它。
At each position $t$, TTT maintains a weight matrix $W_t \in \mathbb{R}^{d \times d}$ (or a small neural network parameterized by $W_t$). The hidden state of the sequence model is $W_t$.在每个位置 $t$,TTT 维护一个权重矩阵 $W_t \in \mathbb{R}^{d \times d}$(或者一个以 $W_t$ 为参数的小网络)。序列模型的隐状态就是 $W_t$。
For each incoming token $x_t \in \mathbb{R}^d$, TTT constructs a self-supervised task:对每个进来的 token $x_t \in \mathbb{R}^d$,TTT 构造一个自监督任务:
where $k_t = x_t W_K \in \mathbb{R}^d$ is the key and $v_t = x_t W_V \in \mathbb{R}^d$ is the target value, with $W_K, W_V \in \mathbb{R}^{d \times d}$ being outer (slow) weights shared across all positions.其中 $k_t = x_t W_K \in \mathbb{R}^d$ 是 key,$v_t = x_t W_V \in \mathbb{R}^d$ 是目标 value,而 $W_K, W_V \in \mathbb{R}^{d \times d}$ 是所有位置共享的外层(慢)权重。
This is a simple ridge regression objective: the fast-weight matrix $W_t$ is asked to map the current key to the current value.这就是一个简单的岭回归目标:要求快权重矩阵 $W_t$ 把当前的 key 映射到当前的 value。
The weight is updated by a gradient step:权重用一步梯度更新:
where $\eta > 0$ is a learned step size (scalar or per-parameter), and the gradient $\nabla_W \mathcal{L}_t(W) = (Wk_t - v_t)k_t^\top \in \mathbb{R}^{d \times d}$ is the outer product of the residual and the key.其中 $\eta > 0$ 是学出来的步长(标量或逐参数),梯度 $\nabla_W \mathcal{L}_t(W) = (Wk_t - v_t)k_t^\top \in \mathbb{R}^{d \times d}$ 是残差与 key 的外积。
The output at position $t$ is then read out by querying the updated model:位置 $t$ 的输出,就是拿 query 去问更新后的模型:
$$z_t = W_t \cdot q_t \qquad z_t \in \mathbb{R}^d$$where $q_t = x_t W_Q \in \mathbb{R}^d$ is the query, again using a shared slow weight $W_Q \in \mathbb{R}^{d \times d}$.其中 $q_t = x_t W_Q \in \mathbb{R}^d$ 是 query,同样用共享的慢权重 $W_Q \in \mathbb{R}^{d \times d}$ 得到。
After $t$ gradient steps, the fast weight $W_t$ encodes the accumulated gradient information from all positions $1, \ldots, t$:走完 $t$ 步梯度之后,快权重 $W_t$ 编码了位置 $1, \ldots, t$ 上累积的全部梯度信息:
$$W_t = W_0 - \eta \sum_{s=1}^{t} \left(W_{s-1} k_s - v_s\right) k_s^\top$$Each $W_t$ is distinct — it depends on the entire history $(x_1, \ldots, x_t)$. This is precisely the per-position weight matrix $W^{(t)}$ from the general model, but expressed implicitly through gradient accumulation rather than stored explicitly.每个 $W_t$ 都不一样——它依赖于整段历史 $(x_1, \ldots, x_t)$。这正是一般模型里的逐位置权重矩阵 $W^{(t)}$,只不过是用梯度累积隐式表达出来的,而不是显式存下来的。
The standard Transformer loops the same static function $f_W$ over positions. TTT loops a function whose weights change at each step via gradient descent. The loop body is not $f_W$ but $f_{W_t}$, and $W_t$ evolves by a local optimization step at each position.标准 Transformer 在各个位置上循环同一个静态函数 $f_W$。TTT 循环的则是一个权重会变的函数,每步都由梯度下降更新。循环体不是 $f_W$ 而是 $f_{W_t}$,而 $W_t$ 在每个位置上都经过一次局部优化。
This gives TTT the expressivity of per-position weights without the $O(T \cdot d^2)$ storage cost. The trade-off: the per-position weight is determined by a fixed optimization trajectory (gradient descent from $W_0$), not a freely learned mapping.这让 TTT 拿到了逐位置权重的表达力,却不必付出 $O(T \cdot d^2)$ 的存储代价。代价在于:逐位置的权重由一条固定的优化轨迹决定(从 $W_0$ 出发的梯度下降),而不是一个可以自由学习的映射。
The TTT update is more than vanilla gradient descent — it is closely related to LocoProp [Amid et al., 2022], a local propagation algorithm for training neural networks layer by layer.TTT 的更新不止是普通梯度下降——它和 LocoProp [Amid et al., 2022] 关系密切,后者是一种逐层训练神经网络的局部传播算法。
LocoProp decomposes the global training objective into local per-layer targets. Each layer is trained to minimize:LocoProp 把全局训练目标拆成每层的局部目标。每一层要最小化的是:
where $h_{\ell-1} \in \mathbb{R}^d$ is the input to layer $\ell$, $\hat{h}_\ell \in \mathbb{R}^d$ is the local target (a detached signal from the next layer's gradient), $W_\ell^{\text{prev}}$ is the previous iterate, and $\|\cdot\|_F$ denotes the Frobenius norm.其中 $h_{\ell-1} \in \mathbb{R}^d$ 是第 $\ell$ 层的输入,$\hat{h}_\ell \in \mathbb{R}^d$ 是局部目标(来自下一层梯度、已 detach 的信号),$W_\ell^{\text{prev}}$ 是上一次迭代的权重,$\|\cdot\|_F$ 表示 Frobenius 范数。
TTT's self-supervised loss $\mathcal{L}_t(W) = \|W k_t - v_t\|^2$ is structurally identical, with the key $k_t$ playing the role of the layer input and the value $v_t$ as the local target. TTT is therefore LocoProp applied horizontally across time rather than vertically across depth.TTT 的自监督损失 $\mathcal{L}_t(W) = \|W k_t - v_t\|^2$ 在结构上与之完全一致:key $k_t$ 扮演层输入的角色,value $v_t$ 充当局部目标。所以 TTT 就是把 LocoProp 横着沿时间轴用了一遍,而不是纵向沿深度用。
Axis: depth (layer $\ell$)轴:深度(第 $\ell$ 层)
Target: signal from adjacent layer目标:来自相邻层的信号
Update: local regression per layer更新:每层一次局部回归
Purpose: credit assignment without full backprop目的:不做完整反向传播的信用分配
Axis: time (position $t$)轴:时间(位置 $t$)
Target: value $v_t$ at current position目标:当前位置的 value $v_t$
Update: local regression per position更新:每个位置一次局部回归
Purpose: per-position weight adaptation目的:逐位置的权重自适应
We can now arrange these architectures on a spectrum of how much weight sharing they enforce along the time axis:现在可以把这些架构按「沿时间轴共享多少权重」排成一个谱:
Full sharing (Standard Transformer): $W^{(t)} = W$ for all $t \in \{1,\ldots,T\}$. One matrix, reused everywhere. Cheapest, but position-invariant.完全共享(标准 Transformer):对所有 $t \in \{1,\ldots,T\}$ 都有 $W^{(t)} = W$。一个矩阵,到处复用。最省,但位置不变。
Gradient-accumulated sharing (TTT): $W^{(t)} = W_0 - \eta\sum_{s \leq t} g_s$, where $g_s$ is the gradient at step $s$. Position-dependent, but constrained to a gradient descent trajectory. $O(d^2)$ hidden state.梯度累积式共享(TTT):$W^{(t)} = W_0 - \eta\sum_{s \leq t} g_s$,其中 $g_s$ 是第 $s$ 步的梯度。依位置而变,但被限制在一条梯度下降轨迹上。隐状态规模 $O(d^2)$。
No sharing (Per-Position Model): $W^{(t)}$ freely chosen for each $t$. Most expressive, but $O(T \cdot d^2)$ parameters — infeasible for long sequences.完全不共享(逐位置模型):每个 $t$ 的 $W^{(t)}$ 自由取值。表达力最强,但参数量 $O(T \cdot d^2)$——长序列下不可行。
The standard Transformer's position-invariance is its defining limitation at long context. When processing position $t = 1{,}000{,}000$, the model uses the exact same weight matrices as at position $t = 1$. The only differentiation comes from the attention pattern over context, not from the processing function itself.位置不变性,是标准 Transformer 在长上下文下最本质的局限。处理位置 $t = 1{,}000{,}000$ 时,模型用的还是位置 $t = 1$ 那组权重矩阵。唯一的区别来自注意力在上下文上的分布,而不是处理函数本身。
TTT breaks this invariance in a principled, memory-efficient way. The fast weight $W_t$ at position $t$ encodes a compressed summary of all keys and values seen so far, updated via a gradient step that costs only $O(d^2)$ per position — the same asymptotic cost as a standard attention operation.TTT 以一种有原则且省显存的方式打破了这种不变性。位置 $t$ 上的快权重 $W_t$ 压缩地记录了此前见过的所有 key 与 value,每个位置只需一步 $O(d^2)$ 的梯度更新——和标准注意力操作是同一个渐近量级。
@misc{deng2026looped,
author = {Chunyuan Deng},
title = {Transformers Are (Naively) Looped Transformers, Horizontally},
year = {2026},
url = {https://charlesdddd.github.io/blog/transformers-are-looped.html}
}