A small habit that helps me think more honestly about new ideas一个帮我更诚实地看待新想法的小习惯
I want to share a small thinking tool I've been playing with. I'm definitely not an established researcher—just someone trying to get better at evaluating ideas honestly. I've found this little habit helpful for myself, so I figured I'd write it up. Maybe it resonates with you, maybe not. Either way, here it is.想分享一个我最近一直在用的小思维工具。我肯定算不上什么资深研究者——只是一个想把「诚实地评估想法」这件事做得更好一点的人。这个小习惯对我自己挺有用,就写下来了。你可能有共鸣,也可能没有,反正先放在这儿。
I think most of us, myself included, have a soft spot for new stuff. There's something about a shiny new idea that just feels better than the thing we already have. In research, I notice this a lot. The community has method A and method B, and someone proposes method C. You read the paper and think, "oh, this is interesting, this is fresh." But sometimes I catch myself realizing that it might not actually be better—it just feels that way because it's new.我觉得大多数人,包括我自己,都对新东西有偏爱。一个闪亮的新想法,感觉上就是比手上已有的东西更好。做研究时我经常注意到这一点:社区里已经有方法 A 和方法 B,有人提出方法 C,你读完论文会想「哦,有意思,挺新的」。但有时我会反应过来,它未必真的更好——只是因为它新,所以感觉更好。
I think novelty and progress can be different things, but at least for me, my brain has a hard time telling them apart sometimes.新颖和进步可以是两回事,但至少对我来说,大脑有时候是分不清的。
So here's the little trick. It's honestly pretty simple, maybe too simple, but it works for me.所以有了这个小技巧。说实话它非常简单,可能简单过头了,但对我有用。
Say the community already has work A and B, and you (or someone) propose C. Instead of evaluating C in the natural order—given A and B, how good is C—try flipping it. Imagine a world where we already had A and C, and someone proposes B instead. Then see which proposal actually feels more compelling.假设社区里已经有工作 A 和 B,你(或者别人)提出了 C。不要按自然顺序去评估 C——「已有 A 和 B,C 有多好」——而是把顺序翻过来:想象一个世界里我们本来就有 A 和 C,这时有人提出 B。然后看看哪一个提案更有说服力。
Formally, compare C | A,B versus B | A,C.写得正式一点,就是比较 C | A,B 与 B | A,C。
That's basically it. The idea is that by swapping the order, you can maybe peel away some of the novelty bias. Instead of comparing "the new thing" against "the established things," you're trying to put them on more equal footing and see what you actually think.基本上就是这样。这么换个顺序,多少能剥掉一层新颖性偏差:你不再是拿「新东西」去比「既有的东西」,而是尽量把它们放在同一起跑线上,看看自己真实的判断是什么。
Let me try walking through an example that's been on my mind. Just to be clear, this is totally my personal take—you might see it differently, and I'd be curious to hear why.举一个最近一直在想的例子。先说清楚,这完全是我个人的看法——你可能有不同意见,我也很想听听为什么。
A—Machines naturally reason in latent space. This isn't really new. In some sense, all of machine learning is about building better representation spaces, finding good features for next-token prediction. The "latent space" has kind of always been there, quietly doing its thing.A——机器本来就在隐空间里推理。这并不新鲜。某种意义上,整个机器学习就是在构造更好的表示空间,为下一个 token 的预测找到好特征。「隐空间」一直都在那儿,安静地干着自己的活。
B—Chain-of-Thought reasoning. Someone proposes that instead of reasoning purely in hidden states, we have models reason over explicit textual tokens. Step by step, in natural language. A pretty cool idea that changed how we think about inference.B——思维链推理。有人提出,与其只在隐状态里推理,不如让模型在显式的文本 token 上推理,用自然语言一步一步来。这个想法很漂亮,也改变了大家对推理阶段的理解。
C—Latent reasoning (the recent wave), along with interpretability work trying to understand what these latent tokens actually encode.C——最近这一波隐式推理,以及试图弄清这些隐式 token 到底编码了什么的可解释性工作。
In the natural order (A,B → propose C): latent reasoning feels exciting. We're going back to continuous representations, moving beyond the constraints of text. And it genuinely is different from text-based CoT, no question.按自然顺序看(A、B → 提出 C):隐式推理让人兴奋。我们回到了连续表示,跳出了文本的约束。它确实和基于文本的 CoT 不一样,这点毫无疑问。
But I keep thinking about what we might be trading away. One of CoT's really nice properties is that humans can read it, write it, supervise it, and scale it. Both humans and machines can generate reasoning traces pretty cheaply. But latent tokens are really hard to write or verify. And it's not obvious to me what we're getting in return.但我一直在想,我们换掉的是什么。CoT 有一个很好的性质:人能读、能写、能监督、能规模化。人和机器都能以很低的成本生成推理轨迹。而隐式 token 很难写,也很难验证。换回来的东西是什么,我并不清楚。
Now try flipping it. Imagine we live in a world with A and C—machines have always had latent reasoning, and people have been working on its interpretability. Now someone walks in and proposes B: "Hey, what if we just make the model reason in plain text?"现在把顺序翻过来。想象我们生活在只有 A 和 C 的世界里——机器一直在隐空间里推理,大家一直在研究它的可解释性。这时有人进来提出 B:「要不,我们干脆让模型用纯文本来推理?」
I think that's a pretty interesting pitch.我觉得这个提案相当有意思。
If you look at the two columns side by side, it seems like C | A,B—latent reasoning, given we already have CoT—is kind of asking us to trade some pretty concrete advantages for something less transparent. Meanwhile B | A,C—CoT, given we already have latent reasoning—feels like a really practical, scalable idea that anyone can inspect and build on.把两栏并排看,C | A,B——在已有 CoT 的情况下提出隐式推理——像是在要求我们拿一些相当具体的优势,去换一个更不透明的东西。而 B | A,C——在已有隐式推理的情况下提出 CoT——则像一个非常实用、可规模化、谁都能检查并在上面继续搭建的想法。
I have my own sense of which proposal feels stronger, and you probably have yours. The point isn't really that one is definitively right—it's more that flipping the order helps me see past the initial reaction and actually think about the tradeoffs more clearly.哪个提案更强,我心里有自己的判断,你大概也有。重点其实不在于哪个一定对——而在于换个顺序,能帮我越过第一反应,把取舍看得更清楚一点。
Here's another one I've been thinking about. The more I work with this stuff, the more I find myself appreciating it.再举一个我一直在想的例子。这东西我用得越多,越觉得它了不起。
The Transformer architecture. I think it's easy to take it for granted—it's been around since 2017, it's everywhere, it kind of fades into the background of modern ML. But when I actually stop and think about what it gives us, it's pretty remarkable.Transformer 架构。它太容易被当成理所当然了——2017 年就有了,到处都是,几乎已经融进现代机器学习的背景里。但真停下来想想它给了我们什么,其实相当惊人。
A lot of architectures are finicky. You spend a lot of time tuning hyperparameters, finding the right configuration, trying to make things competitive. The Transformer just kind of... stands there. (My friend @zhouxiang and I once did a funny cosplay of it—long story.) You train it, and it gives you good performance. It's really quite robust.很多架构都很挑。你要花大量时间调超参、找配置,才能把效果做到能打。Transformer 则是……就那么立在那儿。(我和朋友 @zhouxiang 还一起 cosplay 过它,说来话长。)你把它训起来,它就给你不错的性能。它是真的稳。
And the nice properties keep adding up. Transformers with CoT are theoretically Turing-complete—they can, in principle, compute anything. They maintain exact KV caches, which means fast inference and good recall over long contexts. They seem like a natural fit for post-training, for agentic RL, for a lot of the things people want to build on top.好性质还在不断累加。带 CoT 的 Transformer 在理论上是图灵完备的——原则上什么都能算。它维护精确的 KV cache,因此推理快、长上下文下的召回也好。它天然适配后训练、agentic RL,以及大家想在上面搭的许多东西。
So I tried the ABC test here too. For a new architecture that wants to replace the Transformer: imagine you already had that new architecture (C), and someone came along and proposed the Transformer (B). I think the pitch would sound something like this.所以我在这里也做了一次 ABC 测试。对一个想取代 Transformer 的新架构:想象你手上本来就有那个新架构(C),这时有人过来提出 Transformer(B)。那段推销词大概会是这样。
"Hey, I have this architecture that just works out of the box. It scales predictably. It has an exact memory mechanism. It supports all sorts of post-training. You barely need to tune anything. And paired with textual reasoning, it's theoretically universal."「我有个架构,开箱即用。scaling 行为可预测。带一套精确的记忆机制。各种后训练都支持。几乎不用调什么参。再配上文本推理,理论上还是通用的。」
I don't know about you, but to me that sounds like a really strong proposal. I think in a lot of cases, the Transformer-as-proposal might actually be the more compelling pitch.不知道你怎么想,反正在我听来这是一个很强的提案。很多情况下,「把 Transformer 当作新提案」反而是更有说服力的那一个。
This doesn't mean we should stop exploring new architectures—definitely not. But I think it helps to appreciate the baseline we're comparing against. The Transformer has earned its spot, and I think acknowledging that is just being honest with ourselves.这不是说我们该停止探索新架构——当然不是。但认清我们在跟什么样的 baseline 比,是有好处的。Transformer 这个位置是它自己挣来的,承认这一点,只是对自己诚实而已。
The ABC thing isn't a rigorous framework or anything. It's really just a mental habit—a way to pause before I get too excited about something new and check whether I actually think it's better, or whether it just feels better because it's different.ABC 算不上什么严谨的框架。它只是一个思维习惯——在对一个新东西太上头之前先停一下,问问自己是真觉得它更好,还是仅仅因为它不一样所以感觉更好。
I find it keeps me a bit more honest with myself, especially when I'm evaluating my own ideas. (My own ideas always feel the most novel to me, which is probably exactly when I need this the most.)它让我对自己诚实一点,尤其是在评估我自己的想法时。(我自己的想法在我看来永远最新颖,而这大概正是最需要这一招的时候。)
If you have thoughts on this, or examples where this kind of thinking breaks down, or cases where flipping the order actually reveals the opposite of what you'd expect—I'd genuinely love to hear about them. This is very much a work in progress, like pretty much everything else in my research life.如果你对这个有想法,或者有这种思路失效的例子,又或者有「换个顺序之后结论反而相反」的情况——我都很想听。这东西还远没有成型,就像我研究生活里的其他大部分事情一样。
@misc{deng2026abc,
author = {Chunyuan Deng},
title = {The {ABC} Research Mindset},
year = {2026},
url = {https://charlesdddd.github.io/blog/abc-research-mindset.html}
}