Episode 008

The Pill That Outran Its Explanation

Metformin is taken by well over a hundred million people, and three serious labs now claim three different organs as its true target — a disagreement that is the honest state of the science, not a failure of it.

2026-08-05

Metformin is one of the most-prescribed drugs on Earth and has been in daily use for more than sixty years, yet there is still no settled answer to the simple question of where in the body it does its work. This episode lays out the three rival accounts — the classic liver story, a strong case for the gut, and a 2025 paper from Baylor arguing that a control point in the brain is necessary at low doses in mice. It separates what that new study actually shows, that removing one protein in one brain region silences the drug, from the headline that we have finally cracked metformin. The larger argument is that medicine licenses drugs on whether they work and are safe, not on whether we understand them, so mechanism routinely arrives after use, and sometimes never in full. Along the way it explains why very little of the drug reaching a very sensitive target can still matter, and why not knowing the mechanism carries a real cost.

🎙 Listen · 9:44
Transcript

Follows the audio as it plays — tap any sentence to jump there.

Somewhere near you, probably within a short walk, someone is taking a small white pill with breakfast. It is called metformin.在离你不远的某个地方,也许步行几分钟就到,有人正就着早餐吃下一小片白色药丸。它叫二甲双胍。 On any list of the medicines the human race swallows most often, it would sit near the top.在人类服用最频繁的药物清单上,无论哪一份,它都会排在靠前的位置。 Well over a hundred million people take it, most for type 2 diabetes, some now in the quiet hope that it slows down aging.服用它的人远超一亿,多数是为了治疗 2 型糖尿病,如今也有一些人抱着它能延缓衰老的悄然期望在吃。 It is cheap, and it has been in daily use for more than sixty years. And here is the strange part.它很便宜,而且已经日常使用了六十多年。奇怪的地方就在这里。 Ask a room full of experts exactly how it works, and you will not get one answer. You will get an argument. That is the story tonight.去问一屋子专家它到底是怎么起作用的,你不会得到一个统一的答案。你会得到一场争论。这就是今晚要讲的故事。

Not really a new discovery, more an old and honest confusion that a paper from last year has sharpened rather than settled.这算不上什么新发现,更像是一桩由来已久、坦诚存在的困惑——去年的一篇论文让它变得更尖锐,而非把它解决了。 The headlines said we had finally cracked it.各种头条说我们终于把它弄清楚了。 What actually happened is more interesting, and it tells you something about how medicine really works, which is not the way the textbooks pretend.实际发生的事情要有意思得多,它会让你看到医学真正的运作方式,而那和教科书假装的样子并不一样。

Start with where the drug comes from, because the beginning already contains the lesson. Metformin descends from a plant.先从这种药的来源说起,因为开端本身就已经包含了教训。二甲双胍源自一种植物。 In medieval Europe people grew a herb called goat's rue, also known as French lilac.在中世纪的欧洲,人们种植一种叫山羊豆的草本植物,它也被称为法国紫丁香。 Farmers noticed it made cattle give more milk, and healers gave it to people with the thirst and frequent urination we now recognize as diabetes.农民注意到它能让牛产更多奶,医者则把它给那些有口渴和尿频症状的人——也就是我们如今认识的糖尿病。 In 1914 a French pharmacist pulled the active compound out of the plant. It lowered blood sugar. It was also too toxic to use.1914 年,一位法国药剂师从这种植物中提取出了活性化合物。它能降血糖,但毒性也太大,无法使用。 Chemists in the 1920s built a family of related molecules called biguanides, and one of them, metformin, was tried in people by a French physician named Jean Sterne in 1957, who gave it a hopeful name that translates roughly as glucose eater.20 世纪 20 年代的化学家构建出了一族相关分子,称为双胍类,其中之一就是二甲双胍。1957 年,一位名叫 Jean Sterne 的法国医生在人体上试用了它,并给它取了个满含希望的名字,大致可译为「吃葡萄糖的东西」。 Notice what did not happen in that sequence. Nobody understood the mechanism.注意这一连串过程中没有发生什么。没有人理解它的机制。 They had a plant that worked, then a molecule that worked, and the working came first. Understanding was supposed to catch up later.他们先有了一种有效的植物,再有了一种有效的分子,起效是第一位的。理解本应在之后才跟上。

It has been a long later. And metformin nearly did not survive to see it.这个「之后」等了很久。而二甲双胍差点没能活到那一天。 It had two chemical siblings, phenformin and buformin, sold at the same time.它有两个化学上的同胞——苯乙双胍和丁双胍,当时一同在售。 In the 1970s those two were found to cause a dangerous buildup of acid in the blood, a condition called lactic acidosis, and were pulled from the market in most countries.20 世纪 70 年代,人们发现这两种药会导致血液中危险的酸性物质堆积,一种叫乳酸酸中毒的状况,于是在多数国家被撤出市场。 Metformin causes the same problem, but far more rarely.二甲双胍也会引发同样的问题,但要罕见得多。 Very roughly, phenformin harmed about one patient in four thousand, metformin something closer to one in tens of thousands.非常粗略地说,苯乙双胍大约每四千名患者中伤害一人,二甲双胍则接近每几万人中才有一人。 That gap in the numbers is the whole reason one drug became a global staple and the other two became footnotes.数字上的这道差距,正是一种药成为全球主力、另外两种沦为脚注的全部原因。 It was a matter of degree, and metformin was nearly dragged down with its relatives.这是一个程度的问题,而二甲双胍差一点就被它的亲戚一起拖下水。 Instead it became the drug most doctors reach for first in type 2 diabetes.结果它反而成了多数医生治疗 2 型糖尿病时首先想到的药。

So we have a drug taken by a huge fraction of humanity, with a good record and a clear benefit.所以我们有了这样一种药:全人类中有很大一部分人在服用,记录良好,益处明确。 Now comes the question that should have a simple answer and does not. Where in the body does it actually act?现在轮到那个本该有简单答案、却偏偏没有的问题了。它在体内究竟作用于哪里?

The textbook answer, the one most people carry, is the liver.教科书上的答案,也是大多数人记住的那个,是肝脏。 Your liver makes glucose and releases it into the blood, and the story went that metformin tells the liver to make less, by switching on an energy sensor inside cells called AMPK.你的肝脏制造葡萄糖并把它释放进血液,而那套说法是:二甲双胍通过开启细胞内一种叫 AMPK 的能量感应器,告诉肝脏少造一点。 For years that was the settled picture. Then it started to wobble.多年来这都是既定的图景。然后它开始动摇。 Careful experiments showed you could lower blood sugar with metformin even when you blocked the supposed pathway, which is not what you would expect if that pathway were the whole story.严谨的实验显示,即便你阻断了那条据称的通路,二甲双胍仍能降低血糖——如果那条通路就是全部真相,这可不该发生。

Then a second camp made a strong case for a different organ entirely, the gut. When you swallow metformin, it does not spread evenly.接着第二个阵营为一个完全不同的器官——肠道——提出了有力的论据。当你吞下二甲双胍,它并不会均匀地散布开来。 It piles up in the intestine at concentrations far higher than it ever reaches in the blood. And here is the striking finding.它在肠道中的浓度远高于它在血液中所能达到的水平。而这里有一个引人注目的发现。 You can deliver metformin so that it acts only in the gut, without raising its level in the bloodstream at all, and blood sugar still falls.你可以让二甲双胍只在肠道中发挥作用,而完全不提升它在血液中的浓度,血糖仍然会下降。 The gut, in this account, changes how it handles glucose and signals onward to the liver. This camp has a neat bonus argument.按照这种说法,肠道改变了它处理葡萄糖的方式,并向下游的肝脏发出信号。这一派还有一个漂亮的附加论据。 Metformin's most famous side effect is that it upsets the stomach, especially at first.二甲双胍最出名的副作用是它会引起胃部不适,尤其是在刚开始服用时。 If the gut is where the drug is really working, that side effect is not a random nuisance. It is sitting right next to the mechanism.如果肠道才是这种药物真正起作用的地方,那么这个副作用就不是随机的麻烦。它恰恰紧挨着作用机制。

And now, last year, a third organ walked into the room.而现在,就在去年,第三个器官走进了这个房间。 A team led by Makoto Fukuda at Baylor College of Medicine published a paper in the journal Science Advances with a blunt title.由贝勒医学院的 Makoto Fukuda 领导的一个团队在《Science Advances》期刊上发表了一篇标题直白的论文。 Low dose metformin requires brain Rap1 for its antidiabetic action. Their claim is that the brain is not a bystander but a control room.低剂量二甲双胍的抗糖尿病作用需要脑内 Rap1。他们的论点是,大脑不是旁观者,而是控制室。 Deep in the brain sits a region called the ventromedial hypothalamus, a place that reads the body's fuel state and issues orders about it.大脑深处有一个叫做腹内侧下丘脑的区域,它读取身体的燃料状态并据此发出指令。 In that region are cells carrying a small protein called Rap1.在那个区域里有一些细胞,携带一种叫做 Rap1 的小蛋白。 When metformin reaches those cells, the team found, they fire, and blood sugar drops.该团队发现,当二甲双胍到达这些细胞时,它们就会激活,血糖随之下降。

Two of their experiments are worth picturing, because they are cleaner than most of biology. First, the dose.他们的两个实验值得想象一下,因为它们比大多数生物学实验都更干净利落。首先是剂量。 They injected metformin directly into the brains of mice, and it lowered blood sugar at amounts thousands of times smaller than a normal swallowed dose.他们把二甲双胍直接注入小鼠的大脑,结果它在仅为正常口服剂量千分之几的用量下就降低了血糖。 A few millionths of a gram did it.几百万分之一克就做到了。 Second, and this is the load-bearing one, they used mice engineered to lack that one protein, Rap1, in that one brain region.第二个实验,也是承重的那一个,他们使用了经过基因改造、在那一个脑区缺失那一种蛋白 Rap1 的小鼠。 In those mice, low dose metformin simply stopped working. Their blood sugar did not budge.在这些小鼠身上,低剂量二甲双胍干脆就不起作用了。它们的血糖纹丝不动。 But insulin still worked in them, and so did another diabetes drug. So the animals were not broken in some general way.但胰岛素在它们身上仍然有效,另一种糖尿病药物也有效。所以这些动物并不是以某种笼统的方式坏掉了。 One specific pathway had been switched off, and with it went the effect of one specific drug.有一条特定的通路被关闭了,随之消失的是一种特定药物的效果。 That is how you show a part is not merely present but necessary. You remove it and see what falls silent.这就是你如何证明一个部件不仅仅是存在,而是必需的。你把它移除,看看什么会随之沉默。

Now, there is an obvious objection, and the honest thing is to meet it head on. Metformin is famous for barely getting into the brain.现在,有一个显而易见的反驳,而诚实的做法是正面迎接它。二甲双胍以几乎进不了大脑而闻名。 So how can the brain be where it matters? The answer clears up a confusion that trips people up all the time.那么大脑怎么可能是它起作用的地方呢?答案澄清了一个长期困扰人们的误解。 Quantity is not the same as sensitivity. The brain does not need much metformin because it responds to tiny amounts.数量和敏感性不是一回事。大脑不需要多少二甲双胍,因为它对极微量就有反应。 Think of the thermostat on your wall. It is a small sensor that draws almost no power, and it controls the heating of an entire house.想想你墙上的恒温器。它是一个几乎不耗电的小传感器,却控制着整栋房子的供暖。 A little metformin reaching a very sensitive control point can, in principle, move the whole system.一点点二甲双胍到达一个非常敏感的控制点,原则上就能撬动整个系统。 So the two facts, little reaches the brain and the brain matters, can both be true at once. So what should we take from the paper?所以这两个事实——到达大脑的量很少,而大脑很重要——可以同时成立。那么我们该从这篇论文中得出什么呢?

Here is where the headlines oversold. This is a study in mice. One protein, in one brain region, at low doses.这就是那些标题夸大其词的地方。这是一项在小鼠身上做的研究。一种蛋白,一个脑区,低剂量。 The authors themselves are careful to say they are not ruling out effects elsewhere in the body at higher doses.作者自己也很谨慎地表示,他们并不排除在更高剂量下,药物在身体其他部位也有作用。 And notice what the result does not do. It does not prove the liver camp wrong, or the gut camp wrong.而且请注意这个结果没有做到什么。它并没有证明肝脏派是错的,或肠道派是错的。 It adds a third serious contender to a fight already going. The fair sentence is not we finally know how metformin works.它给一场本就在进行的争论增添了第三个严肃的竞争者。公允的说法不是我们终于知道二甲双胍是如何起作用的。 It is that at low, clinically relevant doses in mice, the brain appears to be necessary. That is a real and interesting claim.而是在小鼠身上、在低的、临床相关的剂量下,大脑看起来是必需的。这是一个真实而有趣的论断。 It is also much smaller than the headline. Step back and the deeper point comes into view.它也比标题所暗示的要小得多。退一步看,更深层的问题才浮现出来。

We have given this drug to hundreds of millions of people for more than sixty years, and there are now at least three respectable answers to where it primarily acts, each defended by a serious laboratory, each calling itself the main event.六十多年来,我们已经把这种药给了数以亿计的人,而关于它主要作用于何处,如今至少有三种站得住脚的答案,每一种都有一家严肃的实验室为之辩护,每一种都自称是主角。 It is tempting to read that as a failure. It is not. It is a fact about medicine that we mostly hide.人们很容易把这解读为一种失败。但它不是。这是医学的一个事实,只是我们大多把它藏了起来。 Drugs are approved on the question does it work and is it safe, not on the question do we understand it.药物获批依据的是它是否有效、是否安全,而不是我们是否理解它。 Aspirin was sold for about seventy years before anyone worked out what it does at the molecular level.阿司匹林卖了大约七十年,才有人搞清楚它在分子层面到底做了什么。 Lithium steadies mood, and we still argue about why. Understanding arrives after use, sometimes long after.锂能稳定情绪,而我们至今仍在争论其原因。理解总是在使用之后才到来,有时要晚很久。

That is not a scandal, but it does carry a cost.这不是什么丑闻,但它确实是有代价的。 When you do not know where a drug acts, you cannot easily predict who it will fail in, or design a cleaner version that hits only the useful target and spares the rest.当你不知道一种药作用于何处时,你就很难预测它会在谁身上失效,也很难设计出一个更干净的版本——只命中有用的靶点,放过其余。 It also matters for the newest hope pinned on metformin, that it might slow aging, now in a large trial.这对于寄托在二甲双胍身上的最新希望同样重要——它或许能减缓衰老,如今正在一项大型试验中接受检验。 Bet on an effect you cannot explain and you are betting blind.押注于一个你无法解释的效应,就是在盲目下注。 Knowing whether the true lever is in the liver, the gut, or the brain is not academic tidiness.弄清真正的杠杆在肝脏、在肠道,还是在大脑,并不是学术上的洁癖。 It is the difference between a drug we inherited and one we can improve.这是一种我们继承下来的药和一种我们能够改进的药之间的区别。

One last thought, for anyone who builds models of complicated systems. The knockout mouse is an ablation.最后一点想法,献给所有为复杂系统构建模型的人。基因敲除小鼠就是一种消融(ablation)实验。 You take a working system, remove one component, and watch whether performance collapses. If it does, that part was carrying weight.你拿一个正常运转的系统,移除其中一个组件,观察性能是否崩溃。如果崩溃了,那个部件就是在承担重量的。 If it does not, it was along for the ride.如果没有,那它不过是搭了个便车。 Same logic whether the system is a mouse or a piece of software, and one of the few clean ways to tell a part that matters from a part that merely happens to be there.无论这个系统是一只小鼠还是一段软件,逻辑都一样——这也是为数不多的、能干净利落地区分一个真正重要的部件与一个只是碰巧存在的部件的方法之一。

So the corrected, one sentence version. Metformin is not a mystery solved.所以,修正后的一句话版本是:二甲双胍并不是一个已经解开的谜。 It is a drug that worked first and is still explaining itself, and the disagreement about where it acts is not the sound of science failing.它是一种先起了作用、至今仍在为自己作解释的药,而关于它作用于何处的分歧,并不是科学失败的声音。 It is the sound of science in the middle of the question.那是科学身处问题之中的声音。

Check your understanding

Try answering before revealing — these are the points the episode turns on.

1. The brain team took mice and deleted a single protein, Rap1, in one brain region, and then low-dose metformin no longer lowered blood sugar. Why is that stronger evidence than simply showing that metformin activates those brain cells?

Showing that metformin lights up some cells only tells you the drug touches them; lots of things are touched without being load-bearing. Deleting the protein tests necessity instead of mere presence. If you remove one component and the drug's effect vanishes, that component was carrying the effect, not just going along for the ride. The design gets even more convincing because insulin and another diabetes drug still worked in the same altered mice, so the animals were not broken in some blanket way. One specific pathway was switched off, and only the drug that depends on it went silent. That is the logic of an ablation, and it is one of the few clean ways to separate a part that matters from a part that merely happens to be there.

2. Metformin is famous for barely crossing into the brain, and it piles up instead in the gut. How can the brain still be a main site of its action without contradicting that fact?

Because how much of a substance reaches a place is a different question from how sensitive that place is. The brain does not need a large dose if it responds to a tiny one, and in the mice a few millionths of a gram injected directly was enough, thousands of times less than a swallowed dose. Think of a thermostat: a small, low-power sensor that governs the heating of a whole house. So little metformin reaching the brain and the brain mattering can both be true at once. The mistake is assuming the organ that holds the most drug must be the organ where the important work happens.

3. There are now three respectable answers to where metformin acts — liver, gut, brain — each from a serious lab. Why should we read that as normal science rather than as a failure?

Because drugs are approved on whether they work and are safe, not on whether we understand them. Efficacy and safety are tested directly in trials; mechanism is a separate scientific question that often gets answered later, sometimes decades later, sometimes never in full. Aspirin was sold for about seventy years before anyone worked out what it does at the molecular level. So a live disagreement about the mechanism of a drug that plainly works is exactly what an unfinished scientific question looks like from the inside. The three camps are not evidence the field failed; they are evidence the field is still in the middle of the question, with each lab having found a real effect and each overstating how central its own effect is.

4. The headlines said we finally know how metformin works. What is the fair version of the claim the Baylor study licenses, and how do the two differ?

The fair version is narrower on several fronts. It is a study in mice, about one protein in one brain region, at low doses, and the authors themselves say they are not ruling out effects elsewhere in the body at higher doses. So the honest sentence is that at low, clinically relevant doses in mice, the brain pathway appears to be necessary. Crucially, that does not prove the liver or gut accounts wrong; it adds a third contender to an argument already underway. The gap between the headline and the claim is the difference between resolving a question and enriching it. Overstating it would repeat the pattern the episode warns about, mistaking a real but partial finding for the whole story.

5. The gut camp treats metformin's stomach side effects as a clue rather than a nuisance. Why would a side effect count as evidence about mechanism?

Because side effects tend to appear where a drug is concentrated and active. Metformin builds up in the intestine at levels far above what it reaches in the blood, and its most common early side effects are digestive. If the gut is genuinely a primary site of action, then the nausea and diarrhea are not random collateral damage happening far from the real work; they are happening right next to it. That co-location is a soft argument, not proof, but it is the kind of consistency you would expect if the mechanism and the side effect share an address. It also shows how the same fact, the drug loves the gut, feeds both a mechanistic claim and a clinical annoyance.

6. If metformin works whether or not we understand it, why does pinning down its mechanism matter in practice?

Because not knowing where a drug acts limits what you can do with it. You cannot easily predict who it will fail in, and you cannot confidently design a cleaner version that hits only the useful target and spares the tissues that produce side effects. It matters especially for the newest hope pinned on metformin, that it might slow aging, now being tested in a large trial; betting on an effect you cannot explain is betting with less information. Knowing whether the true lever sits in the liver, the gut, or the brain is the difference between a drug we inherited and stumbled into using well, and a drug we understand well enough to deliberately improve.

Further reading

  1. Low-dose metformin requires brain Rap1 for its antidiabetic action (Science Advances)Free, open access, technical. The primary paper from the Fukuda lab. This is where the knockout-mouse result and the microgram brain-injection doses come from; read the discussion for the authors' own caution about not excluding effects elsewhere at higher doses.
  2. After 60 Years, Diabetes Drug Revealed to Unexpectedly Affect The Brain (ScienceAlert)Free. A clear plain-language account of the brain finding. Useful, but note it leans toward the finally-solved framing the episode pushes back on.
  3. Metformin's blood sugar control starts in the brain, not just the liver, study finds (News-Medical)Free. Good on the specifics: the ventromedial hypothalamus, the SF1 neurons, and the dose comparison, and it quotes the study's own stated limitation.
  4. Metabolic regulation by the intestinal metformin-AMPK axis (Nature Communications)Free, open access, technical. Represents the gut camp. Read this to see that the brain paper is entering a live argument, not filling a vacuum: the intestine has its own strong claim to be a primary site of action.
  5. Metformin: historical overview (Diabetologia, Bailey 2017)Abstract free; full text may be paywalled depending on access. The reliable source for the goat's rue origin, Jean Sterne in 1957, and the withdrawal of phenformin and buformin for lactic acidosis. Verify the historical dates here.
  6. Metformin: History and mechanism of action (LGC Standards)Free. A short, readable summary of the plant origin, the biguanide family, and the frank admission that the precise molecular mechanism remains unclear after decades of use.
Episode 007

The Fold That Was Always There

An 87-year-old problem fell to three lines of arithmetic, and the reason we can trust the answer is the same reason it stayed a graveyard so long.

2026-08-04

For eighty-seven years the Jacobian conjecture was assumed true, and a long line of mathematicians, some of them famous, published proofs that later collapsed under a single hidden error; it even helped wreck the early career of Yitang Zhang. On the twentieth of July 2026 a mathematician at Anthropic, Levent Alpöge, used the AI system Claude Fable 5 to search out a counterexample in three dimensions, short enough to fit in one social-media post, and the conjecture was refuted overnight. This episode separates what the popular framing gets wrong, that a machine out-reasoned the mathematicians, from what actually happened, a fast unreasoning search aimed by a human and verified by arithmetic. The argument that runs through it is the asymmetry between finding and checking: a proof can hide a fatal error for decades, while a counterexample carries its own verification with it. That single asymmetry explains both why the problem resisted so long and why it fell so fast.

🎙 Listen · 9:45
Transcript

Follows the audio as it plays — tap any sentence to jump there.

Here is a claim that should not be possible.有这样一个说法,它本不该成立。 A problem that had stood for eighty-seven years, that swallowed careers and broke at least one promising one, that famous mathematicians announced they had solved and were then quietly shown to be wrong, was settled last month by a formula short enough to fit in a single post on social media.一个悬置了 87 年的问题,它吞噬了许多人的学术生涯,至少毁掉了一段本来大有前途的职业道路,一些著名数学家宣布自己解决了它,随后又被悄悄地证明是错的——上个月,这个问题被一个短到足以放进一条社交媒体帖子里的公式解决了。 Three lines. And once those three lines were written down, anyone with a pen and an afternoon could check that they were right.三行字。而一旦这三行字被写下来,任何人只要有一支笔、有一个下午的时间,就能核验它们是对的。 I want to spend these ten minutes on why that combination, decades of failure and then a refutation a student could verify, is not a paradox but the whole lesson.我想用这十分钟来谈谈,为什么这种组合——几十年的失败,然后是一个学生就能核验的反驳——不是一个悖论,而恰恰是全部的教益所在。

The problem is called the Jacobian conjecture, and I can state it to you without much machinery.这个问题叫做雅可比猜想(Jacobian conjecture),我可以不借助太多工具就把它讲给你听。 Think of a rule that takes a point in space and moves it to another point, where the rule is built only out of polynomials. 设想有这样一条规则,它把空间里的一个点移动到另一个点,而这条规则完全由多项式构成。 Adding, multiplying, raising to powers, nothing fancier.加法、乘法、乘幂,没有比这更花哨的东西。 Now there is a quantity you can compute for such a rule, called the Jacobian determinant, which measures, at each point, how the rule is stretching and twisting space right there.现在,对这样一条规则,你可以计算出一个量,叫做雅可比行列式(Jacobian determinant),它在每一点上度量这条规则此处正如何拉伸和扭曲空间。 If that quantity is never zero, the rule does not crush any little patch of space down to nothing.如果这个量从不为零,那么这条规则就不会把空间中的任何一小块压缩成无。 It stays, in the language of the subject, locally invertible. Everywhere you look up close, it can be undone.用这门学科的语言来说,它保持局部可逆。你就近处处观察,它都能被还原。

The question is whether local always adds up to global.问题在于,局部是否总能累加成整体。 If the rule never crushes anything anywhere, and in fact its Jacobian is a nonzero constant, the same value everywhere, must the rule be undoable as a whole?如果这条规则处处都不压缩任何东西,而且它的雅可比行列式实际上是一个非零常数,处处取同一个值,那么这条规则作为整体是否就一定可以被还原? Must there be a second polynomial rule that perfectly reverses it, sending every point back where it came from?是否一定存在第二条多项式规则,能完美地把它逆转过来,把每个点都送回它原来的地方? For eighty-seven years the answer everyone expected was yes. It feels like it has to be yes.87 年来,所有人预期的答案都是肯定的。感觉上它必然是肯定的。 If you can undo the map in every neighbourhood, surely you can undo it everywhere.如果你能在每一个邻域里还原这个映射,那么你当然应该能处处将它还原。

But there is a gap between local and global, and it is easy to picture. Imagine wrapping a long strip of paper around and around a cylinder.但局部与整体之间存在一道缝隙,而且很容易想象。设想把一长条纸绕着一个圆柱一圈又一圈地缠上去。 At every point the strip lies flat and neat, and locally nothing is wrong.在每一点上,纸条都平整服帖,局部看没有任何问题。 But globally the strip overlaps itself, so a single point on the cylinder has several layers of paper stacked above it.但从整体上看,纸条与自身重叠,于是圆柱上的某一个点上方就叠着好几层纸。 Locally one to one, globally many to one.局部一一对应,整体多对一。 The Jacobian conjecture was the bet that for polynomial rules with a constant nonzero Jacobian this could never happen.雅可比猜想赌的就是:对于雅可比行列式为非零常数的多项式规则,这种情况永远不会发生。 That the map could never fold space back onto itself.赌这个映射永远不可能把空间折叠回它自身之上。 Keller wrote the conjecture down in nineteen thirty-nine, and some trace the question further back, to the eighteen eighties.Keller 在 1939 年写下了这个猜想,还有人把这个问题追溯得更远,一直追到 1880 年代。

Now, the thing you need to feel about this problem is how it treated the people who attacked it.现在,关于这个问题你需要体会的一点,是它如何对待那些攻克它的人。 It is one of the most notorious graveyards in mathematics.它是数学中最声名狼藉的坟场之一。 Over the decades there were many published proofs, and then the proofs were read closely, and a subtle error was found in each, and the proof collapsed.几十年间,有许多已发表的证明,然后这些证明被仔细研读,人们在每一个之中都发现了一处微妙的错误,于是证明便崩塌了。 This happened to serious people.这发生在一些严肃认真的人身上。 Beniamino Segre and Wolfgang Gröbner, two of the significant mathematicians of the twentieth century, each put forward arguments that did not survive.Beniamino Segre 和 Wolfgang Gröbner,二十世纪两位举足轻重的数学家,各自提出了论证,却都没能站住脚。 The encyclopedias eventually added something close to a health warning to the entry, telling readers that any new proof of the Jacobian conjecture should be presumed wrong until shown otherwise.百科全书最终为这个词条添上了近乎一则健康警告的东西,告诉读者:任何关于雅可比猜想的新证明,在被证明成立之前,都应被推定为错的。

And it did real damage to at least one life.而它确实对至少一个人的一生造成了真实的伤害。 In nineteen ninety-one a man named Yitang Zhang finished his doctorate at Purdue, on this exact conjecture.1991 年,一个名叫张益唐的人在普渡大学完成了他的博士学位,做的正是这个猜想。 His thesis leaned on a result his own advisor had published, and that result later turned out to be flawed.他的论文依赖于他导师本人发表过的一个结果,而那个结果后来被证明是有缺陷的。 Zhang's work never got published, he and his advisor fell out, and he left without the recommendation letters an academic career needs.张的工作从未得以发表,他和导师闹翻,离开时也没有拿到学术生涯所需的推荐信。 For years afterward he drifted through ordinary jobs, at one point working behind the counter of a Subway sandwich shop, a trained mathematician with no position.此后多年,他辗转于普通的工作之间,一度在一家 Subway 三明治店的柜台后打工——一个受过训练的数学家,却没有职位。 It was not until he was in his late fifties that Zhang did something extraordinary, proving a landmark result about the gaps between prime numbers and becoming, very late, famous.直到年近六十,张才做出了一件非同寻常的事:证明了一个关于素数间隔的里程碑式结果,并在很晚的时候成了名。 But the years before that he lost partly to this conjecture. He is the human cost of a problem that looked simple and was not.但在那之前的那些年,他有一部分是输给了这个猜想。他就是一个看似简单、实则不然的问题所付出的人的代价。

So that is the board.这就是全局。 Eighty-seven years, a pile of dead proofs, a warning label, a wounded career, and a near-universal belief that the statement was true and just needed the right argument.87 年,一堆夭折的证明,一个警示标签,一段受创的职业生涯,以及一种近乎普遍的信念——认为这个命题是真的,只是需要找到正确的论证。 And then, on the twentieth of July this year, a mathematician named Levent Alpöge, who works at the company Anthropic, posted a counterexample.然后,就在今年 7 月 20 日,一位名叫 Levent Alpöge 的数学家——他在 Anthropic 这家公司工作——发布了一个反例。 Not a proof that the conjecture is true.不是一个证明猜想为真的证明。 A specific rule, in three dimensions, whose Jacobian determinant is the constant minus two, never zero, that nonetheless sends three different points to the very same output.而是一条具体的规则,在三维空间中,它的 Jacobian 行列式是常数 −2,永不为零,却仍然把三个不同的点送到了完全相同的输出。 A map that folds.一个会折叠的映射。 If three points land on one, the map cannot be undone, because there is no rule that could know which of the three to send that point back to.如果三个点落到同一个点上,这个映射就无法被逆转,因为没有任何规则能知道该把那个点送回三者中的哪一个。 The conjecture, in three dimensions and above, is simply false. It had been false the whole time.这个猜想,在三维及以上,根本就是假的。它一直都是假的。 There was a fold hiding in the space of polynomials, and for eighty-seven years no one had found it.在多项式的空间里藏着一处折叠,而 87 年来没有人找到它。

Two things about how it was found, and they are the point of the episode. The first is that Alpöge did not find it by hand.关于它是如何被找到的,有两点,也正是本期的重点。第一点是,Alpöge 不是靠手算找到它的。 He used an artificial intelligence system, called Claude Fable 5, to search.他用了一个叫 Claude Fable 5 的人工智能系统来搜索。 And here you have to be careful, because the headlines took this and ran somewhere it should not go.而在这里你必须小心,因为各种标题抓住这一点,把它带到了一个不该去的地方。 They said, in effect, that the machine had solved a problem the mathematicians could not. That is the wrong lesson twice over.它们实际上是在说,机器解决了数学家们无法解决的问题。这个说法在两个层面上都是错的。 It did not solve the problem. It refuted it, which is a different and much cheaper kind of act.它并没有解决这个问题。它反驳了这个问题,而这是一种不同的、成本低得多的行为。 And it did not do the part that mathematicians spend their lives on, the long reasoning, the proving. What it did was search.而且它并没有做数学家们穷尽一生去做的那部分工作——漫长的推理,证明。它所做的是搜索。 The difficulty here was never a deep chain of logic.这里的难点从来都不是一条深奥的逻辑链条。 The difficulty was that the counterexample was one needle in an unimaginably large haystack of possible polynomial rules, and no human had a good way to rummage through that haystack fast enough.难点在于,这个反例是一根针,藏在一个大到无法想象的、由所有可能的多项式规则构成的干草堆里,而没有哪个人有好的办法能足够快地翻遍那个干草堆。 A system that can generate and test candidates at enormous speed is exactly the right tool for a search like that, and exactly the wrong tool to trust for a proof.一个能以极高速度生成并测试候选对象的系统,正是这类搜索所需要的恰当工具,也恰恰是最不该用来托付一个证明的工具。

Which brings me to the second thing, and the real argument.这就引出了第二点,也是真正的论点。 Why should you believe this result, when you should not have believed any of the eighty-seven years of proofs before it?既然此前 87 年里的任何证明你都不该相信,那你为什么应该相信这个结果? Because finding and checking are not the same difficulty.因为寻找和检验并不是同一种难度。 A proof of the conjecture is a long argument, and a long argument can hide a fatal error in a single line, which is precisely how Segre and Gröbner and Zhang's advisor came to grief.对猜想的一个证明是一段漫长的论证,而漫长的论证可能在某一行里藏着一个致命的错误——Segre、Gröbner 以及张的导师,正是这样栽了跟头。 To check a proof you must follow every step and trust every one. But a counterexample is not an argument. It is an object.要检验一个证明,你必须跟随每一步,并信任每一步。但反例不是论证。它是一个对象。 To check Alpöge's counterexample you do not follow any reasoning at all.要检验 Alpöge 的反例,你根本不需要跟随任何推理。 You take his three lines, you compute the Jacobian determinant and confirm it is minus two, and you plug in the three points and confirm they collide.你取他那三行,计算 Jacobian 行列式并确认它是 −2,再代入那三个点并确认它们撞在了一起。 It is arithmetic.这是算术。 A careful student can do it in an afternoon, and in effect thousands of people did, within hours, which is why the result was accepted almost at once even though it has not yet been through formal peer review.一个细心的学生一个下午就能做完,而实际上有成千上万人在几个小时之内就做了,这正是为什么这个结果几乎立刻被接受,尽管它还没有经过正式的同行评审。 The machine's answer did not have to be trusted. It had to be checked, and checking was cheap.机器给出的答案不需要被信任,它只需要被核对,而核对是廉价的。

That asymmetry is the whole story, and it explains both halves of the mystery at once.这种不对称就是整个故事,它一举解释了这个谜题的两半。 It explains why the problem resisted for eighty-seven years, because everyone was trying to build a proof, the hard direction, when the truth lay in the easy direction all along, waiting for someone to search instead of argue.它解释了为什么这个问题顽抗了 87 年,因为所有人都试图去构建一个证明,也就是困难的那个方向,而真相自始至终就躺在容易的那个方向,等着有人去搜索,而不是去论证。 And it explains why the refutation, once it came, was believed overnight, because a counterexample carries its own verification with it.它也解释了为什么这个反驳一旦出现,就在一夜之间被人相信,因为一个反例自带对它自身的验证。 The popular framing, that a mind of some new kind has begun to out-reason the mathematicians, has it backwards.流行的说法,即某种新型的心智已经开始在推理上胜过数学家,把事情说反了。 What happened was the opposite of deep reasoning.所发生的恰恰是深度推理的反面。 It was fast, tireless, unreasoning search, aimed by a human who understood which haystack to point it at, and then confirmed by a community doing the one thing machines still cannot certify for us, which is checking that a claim is really true.那是快速、不知疲倦、不加推理的搜索,由一个懂得该把它指向哪一片草垛的人来瞄准,然后由一个共同体来确认,这个共同体所做的正是机器至今仍无法替我们证实的那一件事,也就是核对一个断言是否真的为真。

Two honest cautions before I stop. The counterexample lives in three dimensions and above.在我收尾之前,有两点诚实的提醒。这个反例存在于三维及以上。 The original problem in the plane, in just two variables, is still open, and it is a genuinely different beast, tied to other hard questions, and no one should assume it will fall the same way.平面上的原始问题,只有两个变量的那个,仍然悬而未决,而且它是一头真正不同的野兽,与其他难题纠缠在一起,没有人应该假定它会以同样的方式倒下。 And the exact record of how the machine was prompted has not been fully made public, so the story of the collaboration is thinner than the mathematics is.而且机器究竟是如何被提示的,其确切记录并未完全公开,所以这场协作的故事比数学本身要单薄。 But the mathematics itself is not in doubt, and it cannot be, for the same reason it was accepted so fast.但数学本身并无疑问,也不可能有疑问,原因和它被如此迅速地接受是同一个。 You do not have to take anyone's word for a fold. You can go and find the three points that land on one, and see it for yourself.对于一个折叠,你不必听信任何人的话。你可以自己去找到落在同一点上的那三个点,亲眼看见它。

So here is the sentence to keep. The Jacobian conjecture was not solved by a machine that out-thought us.所以,这里有一句值得记住的话。雅可比猜想并不是被一台在思考上胜过我们的机器解决的。 It was disproved by a search that found what proof after human proof had wrongly ruled out, and the reason we can trust the answer is the very reason the problem was so cruel in the first place.它是被一次搜索所推翻的,这次搜索找到了一个又一个人类证明错误地排除掉的东西,而我们之所以能信任这个答案,恰恰就是这个问题一开始如此残酷的那个原因。 Proofs are hard to check, and a counterexample checks itself.证明难以核对,而一个反例核对它自己。

Check your understanding

Try answering before revealing — these are the points the episode turns on.

1. Why does a nonzero constant Jacobian guarantee that the map is invertible near any single point, and yet not guarantee it is invertible everywhere at once?

A nonzero Jacobian at a point means the map does not flatten space there, so by the inverse function theorem you can undo it in a small neighbourhood. Making the Jacobian a nonzero constant just says this holds at every point, so the map is locally reversible everywhere. But local reversibility is a statement about small patches viewed one at a time. It says nothing about whether far-apart patches might land on top of each other. Global invertibility is the stronger claim that no two points anywhere share an output, and that extra claim is exactly what the conjecture wrongly assumed followed for free.

2. The episode pictures a paper strip wound around a cylinder. What does that image capture about the counterexample?

On the strip, every little region lies flat and behaves perfectly, so nothing is wrong locally. But because the strip wraps around and overlaps itself, one spot on the cylinder can have several layers stacked above it, so several points of the strip map to the same place. That is the difference between locally one-to-one and globally many-to-one. Alpöge's map does the polynomial version of this: its Jacobian is a healthy constant everywhere, yet it folds so that three distinct inputs arrive at a single output, which is precisely what a truly reversible map can never do.

3. For eighty-seven years the danger in this problem was published proofs that later failed. Why is a counterexample not exposed to that same danger?

A proof is a chain of reasoning, and a single wrong link anywhere breaks the whole thing, which is how careful arguments by Segre, by Gröbner, and by Zhang's advisor all came apart once someone found the flawed step. Verifying a proof means checking every step and trusting each one. A counterexample is not a chain of reasoning at all; it is a concrete object you test directly. You compute its Jacobian and confirm it is the stated constant, then plug in the offending points and watch them collide. There is no long argument that could quietly contain an error, because there is no argument, only arithmetic.

4. In what sense is finding this counterexample a very different kind of task from proving a theorem, and why does that make an AI search a sensible tool here?

Proving a theorem means constructing a valid line of reasoning that no one can break, which is what mathematicians train their judgement to do. Finding this counterexample meant something more like search: somewhere in an enormous space of possible polynomial maps sat one with the right rare properties, and the obstacle was navigating that space quickly rather than reasoning deeply. A system that can generate and test huge numbers of candidates fast is well matched to a needle-in-a-haystack search. It is not thereby trustworthy as a producer of proofs, because for a proof there is no cheap check, only the slow reading of every step.

5. The same asymmetry between finding and checking is used to explain two different things at once. What are they?

First, it explains why the problem held out for eighty-seven years. Everyone was working in the hard direction, trying to build a proof that the conjecture was true, while the actual truth sat in the easy direction, a counterexample waiting to be searched for rather than argued into existence. Second, it explains why the refutation was believed within hours despite not yet being peer reviewed. Because a counterexample can be checked by simple computation, the community could confirm it almost immediately without having to trust the person or the machine that produced it. Cheap checking is what made the result both slow to find and fast to accept.

6. Two cautions temper the story. What are they, and does either put the mathematical result in doubt?

The first caution is that the counterexample only settles dimension three and above; the original conjecture in two variables, the plane case, is still open and is connected to other deep questions, so it should not be assumed to fall the same way. The second is that the precise account of how the AI was prompted has not been fully released, so the collaboration story is thinner than the mathematics. Neither caution touches the core result. The refuting map can be verified by anyone by direct computation, so its correctness does not depend on the undisclosed process or on resolving the still-open planar case.

Further reading

  1. 'hello there the jacobian conjecture is false thanx': why a tiny social media post has mathematicians rethinking AI (The Conversation)Free. The clearest plain-language account of the counterexample and the false-proof history, and honest that the exact AI prompting was not disclosed.
  2. Why a tiny social media post has mathematicians rethinking AI (Phys.org)Free. A near-identical syndication of the piece above; useful if the original is slow to load. Confirms dimension three, Jacobian minus two, and that the two-variable case is still open.
  3. Jacobian Conjecture (Wolfram MathWorld)Free, technical. The precise statement over the complex numbers, why the constant-Jacobian condition is necessary, and the special status of the still-open planar case.
  4. Yitang Zhang (Wikipedia)Free. Background on the doctorate built on his advisor's flawed lemma, the lost years, and the later bounded-gaps-between-primes result. Verify names and dates here before quoting.
  5. Jared Duker Lichtman on the refutation (X post)Free but a social-media post, so treat as a working mathematician's first reaction, not a vetted source. States the inverse-function-theorem framing cleanly.
  6. The AI Revolution in Math Has Arrived (Quanta Magazine)Free. Wider context on where AI is and is not useful in mathematics, helpful for judging how much weight the 'machine out-thought us' framing deserves.
Episode 006

The Planet You Cannot See

The headlines say Venus is tearing itself apart right now; the real result is quieter, cleverer, and honest about what it cannot yet prove.

2026-08-03

You cannot see the surface of Venus, cannot land on it, and cannot put a seismometer there. So how do you tell whether a hidden planet is still geologically alive? A study out of ETH Zurich and Freiburg, published in Nature Geoscience in late July 2026, found an answer in thirty-year-old radar: the shoulders of a rift valley keep time. On Venus's hot crust a raised rift flank slumps and flattens from within, with no weather needed, so a steep flank has to be young. This episode walks through that clock, then separates what the work actually showed from the headline that Venus is 'tearing itself apart right now.' The evidence reaches to recently alive, not to caught in the act, and it matters that the scientists say so.

🎙 Listen · 8:56
Transcript

Follows the audio as it plays — tap any sentence to jump there.

You have never seen the surface of Venus. Neither has anyone.你从未见过金星的表面。谁都没见过。 It is the nearest planet to us, the brightest thing in our sky after the Sun and the Moon, and it is wrapped in a permanent deck of cloud so complete that from outside the whole world looks like a smooth, blank pearl.它是离我们最近的行星,是除太阳和月亮之外我们天空中最亮的东西,而它被一层永久、完整到从外面看整个世界就像一颗光滑、空白的珍珠的云层所包裹。 Under that cloud the air is carbon dioxide, about ninety times heavier than the air you are breathing, and the ground sits at around four hundred and sixty-five degrees, hot enough to melt lead.在那层云之下,空气是二氧化碳,大约比你正在呼吸的空气重 90 倍,而地面温度在 465 度上下,足以熔化铅。 The Soviet Venera landers that reached it in the nineteen seventies and eighties sent back a handful of photographs and then died within an hour or two, cooked and crushed.在 20 世纪七八十年代抵达那里的苏联金星号着陆器传回了寥寥几张照片,随后在一两个小时之内就报废了,被烤熟、被压垮。 So here is the problem I want to spend these ten minutes on.所以这就是我想用这十分钟来谈的问题。 If you cannot see a planet's surface, and you cannot stand on it, how would you ever know whether it is still alive inside?如果你看不到一颗行星的表面,也无法站在它上面,那你究竟怎么才能知道它内部是否还活着?

Because the popular story about Venus is that it is not. You will have heard Venus called Earth's dead twin.因为关于金星的流行说法是它并没有。你多半听过有人把金星称作地球死去的孪生兄弟。 A world that went wrong, boiled dry, seized up, and stopped moving.一个出了岔子的世界,被煮干、卡死、停止了运动。 And at the end of July a paper came out, from a group led at ETH Zurich in Switzerland, that the headlines turned straight into the opposite slogan.而在七月底,一篇论文出炉了,来自瑞士苏黎世联邦理工学院(ETH Zurich)牵头的一个团队,标题党们把它直接变成了相反的口号。 Venus, they announced, is tearing itself apart right now.他们宣布,金星此刻正在把自己撕裂。 I want to argue that both of those slogans are wrong, and that the space between them is the real result, which is more interesting than either.我想主张,这两句口号都是错的,而它们之间的那片空间才是真正的结论,比两者中的任何一个都更有意思。

Start with why anyone ever thought Venus was dead. Unlike Earth, Venus has no plate tectonics.先从为什么会有人认为金星死了说起。与地球不同,金星没有板块构造。 Its outer shell is one continuous piece, not a jigsaw of moving plates, so it does not have our system of grinding boundaries and recycling crust.它的外壳是一整块连续的整体,而不是一副由移动板块拼成的拼图,所以它没有我们那套相互碾磨的边界、循环再生地壳的系统。 That much is genuinely true. But here is the thing that should have stopped the word dead from ever taking hold.这一点是千真万确的。但有件事本该阻止「死了」这个词站住脚。 When Venus was finally mapped, by radar, we could count the craters on it.当金星终于被雷达绘制成图时,我们能数出它上面的撞击坑。 A truly dead world, sitting still for billions of years, accumulates impact craters the way an unswept floor accumulates dust.一个真正死去的世界,静止不动数十亿年,累积撞击坑的方式就像没人打扫的地板积灰。 The Moon is covered in them. Venus is not. Its surface is startlingly clean, which means it is startlingly young.月球上布满了撞击坑。金星没有。它的表面干净得惊人,这意味着它年轻得惊人。 The rock you are looking at is on average only a few hundred million years old, on a planet that is four and a half billion years old.在一颗年龄 45 亿年的行星上,你所看到的岩石平均只有几亿年的历史。 Something has been repaving Venus, wiping the craters away, well within its recent past.在其相当晚近的过去,有某种东西一直在给金星重新铺面,把撞击坑抹去。 So the idea that Venus is inert was never really the finding. The finding, for decades, has been the opposite.所以金星是惰性、无活动的这一想法从来就不是真正的发现。几十年来,发现恰恰相反。 Something down there has been busy.那下面有某种东西一直忙碌着。 The only question was whether it is still busy today, or whether it did its repaving in one great spasm long ago and has been quiet ever since.唯一的问题是它今天是否仍在忙碌,还是在很久以前一次巨大的痉挛中完成了重新铺面,此后便一直归于沉寂。

And that question runs straight into the wall I started with. You cannot see the surface, so you cannot watch it move.而这个问题径直撞上了我开头讲的那堵墙。你看不到表面,所以你无法观察它的运动。 You cannot land a network of seismometers, because your instruments last about an hour. What you actually have is old radar.你无法部署一张地震仪网络,因为你的仪器只能撑大约一个小时。你手上实际拥有的是老旧的雷达数据。 The best map we own came from a NASA probe called Magellan, which orbited Venus in the early nineteen nineties, bounced radar down through the cloud, and read the echoes to build a picture of the shape of the ground.我们所拥有的最好的地图来自一台名叫麦哲伦(Magellan)的 NASA 探测器,它在 20 世纪九十年代初绕金星运行,把雷达波穿过云层反射下去,再读取回波来构建地面形状的图像。 That probe has been dead since nineteen ninety-four. The new result did not come from a new spacecraft.那台探测器自 1994 年起就报废了。这项新成果并非来自一艘新的航天器。 It came from looking again, harder, at thirty-year-old data. Hold on to that, because it is half the point. So what did they look at.它来自对三十年前的数据重新、更用力地再看一遍。记住这一点,因为这是要点的一半。那么他们看了什么。

They looked at rifts. A rift is what you get when a planet's crust is pulled apart.他们看的是裂谷。裂谷就是一颗行星的地壳被拉扯分开时形成的东西。 The ground stretches, thins, and drops down along a line, leaving a long valley. Earth has these too.地面沿着一条线伸展、变薄、下陷,留下一条长长的谷地。地球上也有这些。 The East African Rift, where the continent is slowly splitting, is the classic example. Venus has them on a scale that dwarfs ours.东非大裂谷——那里大陆正在缓慢分裂——就是典型的例子。金星上的裂谷规模让我们的相形见绌。 Some of its rift valleys run for ten thousand kilometres, far enough to wrap most of the way around the planet.它的一些裂谷绵延一万公里,足以绕行这颗行星大半圈。 And here is the feature the study is built on.而这就是这项研究所依托的特征。 When you pull the crust apart and drop a valley down the middle, the two shoulders on either side bow upward.当你把地壳拉开、让中间的谷地陷落下去时,两侧的肩部会向上拱起。 They rise into long, broad ridges running parallel to the valley. Geologists call them rift flanks.它们隆起成平行于谷地延伸的长而宽的山脊。地质学家称之为裂谷侧翼(rift flanks)。 Picture the raised lips on either side of a crack in dried mud. Those raised lips are the flanks.想象干裂泥浆中一道裂缝两侧翘起的边唇。那些翘起的边唇就是侧翼。

Now comes the clever idea, and it is the whole episode, so let me go slowly.现在来到那个巧妙的想法,它就是整期节目的核心,所以让我慢慢讲。 On Earth, once a mountain or a ridge stops being pushed up, what wears it down is weather. Rain, ice, rivers, wind.在地球上,一座山或一道山脊一旦不再被抬升,把它磨蚀掉的是天气。雨、冰、河流、风。 It takes millions of years, and it grinds the high ground flat from the outside. But Venus has no rain and no rivers.这要花上数百万年,从外部把高地磨平。但金星没有雨,也没有河流。 There is nothing to erode those flanks from above. And yet the group's computer models showed that the flanks should still not last.没有任何东西能从上方侵蚀那些侧翼。然而研究团队的计算机模型显示,这些侧翼仍然不该长存。 On Venus the rock itself is the culprit.在金星上,罪魁祸首是岩石本身。 The crust is so hot that over time it behaves less like something rigid and more like something slow and stiff and yielding, closer to cold honey than to stone.地壳如此炽热,以致随着时间推移,它的行为不再像坚硬之物,而更像某种缓慢、黏稠、易于变形的东西,更接近冷蜂蜜而非石头。 A raised flank sitting on hot, soft rock cannot hold itself up. It sags.一道隆起的侧翼坐落在炽热、柔软的岩石之上,无法支撑自身。它会下沉。 It relaxes back down and spreads out, from the inside, without any weather touching it at all. The model gives a clear timetable for this.它松弛下来、向外摊开,是从内部发生的,全程没有任何天气触碰到它。模型给出了一份清晰的时间表。 While the rift is young and still pulling apart, its flanks are tall, steep, and narrow.当裂谷还年轻、仍在拉张时,它的侧翼高耸、陡峭而狭窄。 Once the pulling stops, the flanks slump, flatten, and broaden, and they do it relatively fast in geological terms.一旦拉张停止,侧翼便坍塌、变平、变宽,而且以地质学的标准来看,这发生得相对迅速。 So the shape of the shoulder becomes a clock. A steep, high, sharp-edged flank has to be young. A low, wide, gentle one is old.于是肩部的形状成了一座时钟。陡峭、高耸、棱角分明的侧翼必定年轻。低矮、宽阔、平缓的则古老。

And when they went back to the Magellan data with that clock in hand, they found flanks along some of the largest rift systems that are still tall and still steep.而当他们带着这座时钟重新审视麦哲伦(Magellan)数据时,发现沿着一些最大的裂谷系统的侧翼依然高耸、依然陡峭。 Too steep, the argument runs, to have been sitting quiet for the hundreds of millions of years that people had assumed.按这一论证,太陡了——陡到不可能像人们过去所设想的那样,安静地待了数亿年。 If the flank were that old, on rock that soft, it should have slumped by now. It has not. So the rifts, they conclude, are young.如果侧翼真有那么古老,坐落在如此柔软的岩石上,它到现在早该坍塌了。可它没有。所以他们的结论是,这些裂谷是年轻的。 Either they were pulling apart in the very recent past, or, and this is the strong version, they are still pulling apart today.要么它们是在极晚近的过去发生拉张,要么——这是更强的说法——它们至今仍在拉张。

Now I have to do the thing this podcast tries to do every time, which is to separate what was shown from what was claimed.现在我得做这档播客每次都努力去做的事,也就是把已经证明的与所宣称的区分开来。 What the group actually demonstrated is a mechanism and a match.这个团队实际证明的是一套机制和一处吻合。 They built a physical model of how a rift flank rises and then relaxes, and they showed that the steep flanks in the old radar are hard to explain unless the rifting is geologically recent.他们建立了一个物理模型,说明裂谷侧翼如何隆起、然后又松弛下来,并且表明:除非裂谷作用在地质意义上是晚近的,否则老雷达数据里那些陡峭的侧翼很难解释。 That is a real and careful result. What the headlines did with it is another matter. You will have seen the number.这是一个真实而审慎的结果。至于头条新闻拿它做了什么,则是另一回事。你大概见过那个数字。 Venus, they wrote, is rifting at three to ten centimetres a year, comparable to the speed of Earth's plates.他们写道,金星正以每年3到10厘米的速度裂开,与地球板块的速度相当。 That figure is not a measurement. Nobody clocked Venus moving.这个数字不是一次测量。没有人真的测到金星在移动。 It is a rate that comes out of the model when you ask what speed would produce flanks shaped like the ones we see.它是当你问:什么样的速度会造出我们所看到的这种形状的侧翼时,从模型里得出的一个速率。 And the honest word the scientists themselves use is not now but recent, where recent, in this business, can mean anytime within the last tens of millions of years.而科学家自己所用的诚实措辞,不是「现在」,而是「晚近」,在这一行里,「晚近」可以指过去数千万年内的任何时候。 So when a headline says Venus is tearing itself apart right now, it has quietly turned a careful inference about the recent past into a live broadcast of the present.所以当一条头条说金星此刻正在把自己撕裂时,它悄悄地把一个关于晚近过去的审慎推断,变成了对当下的现场直播。 The evidence does not reach that far. It reaches to recently alive. It does not yet reach to caught in the act.证据没有伸得那么远。它伸到了「近来仍活跃」。它还没伸到「正被当场抓个正着」。

That gap is not a failure of the work.这个空缺并不是这项工作的失败。 It is exactly where the work stands, and saying so plainly is the difference between science and a slogan.它恰恰标示了这项工作目前所处的位置,而坦率地这么说,正是科学与口号之间的区别。 What would close the gap is a real measurement, and for once the instrument is on the way.能弥合这一空缺的是一次真实的测量,而这一次,仪器正在路上。 Europe is building a mission called EnVision, and one of the paper's own authors sits on its team.欧洲正在建造一项名为 EnVision 的任务,而这篇论文的一位作者本人就在它的团队里。 It will carry radar sharp enough to look at these same flanks again, years apart, and see whether anything has actually shifted.它将携带足够清晰的雷达,隔上数年再次观测这些同样的斜坡,看看是否真的有什么发生了移动。 Then the inference becomes a fact, or it does not.到那时,这个推断要么变成事实,要么不然。 Until it flies, the correct statement is that we have a good physical reason to think Venus is still moving, not a photograph of it moving.在它升空之前,正确的说法是:我们有充分的物理理由认为金星仍在活动,而不是拥有一张它正在活动的照片。

So let me give you the corrected sentence, the one to keep. Venus was never shown to be dead.所以让我给你一句修正过的话,一句值得记住的话。金星从未被证明是死寂的。 Its clean young face said the opposite all along. And it has not now been shown to be tearing itself apart before our eyes.它那洁净而年轻的面貌一直在说着相反的事。而如今,它也并没有被证明正在我们眼前把自己撕裂。 What actually happened is quieter and better than both.真正发生的事情比这两者都更安静,也更美好。 Someone realised that the shoulders of a rift keep a kind of time, that a steep slope on hot rock is a young slope, and they read that clock in thirty-year-old echoes from a spacecraft that has been silent for three decades.有人意识到,一道裂谷的两肩记录着某种时间,热岩之上的陡坡是年轻的斜坡,于是他们从一艘沉默了三十年的航天器所留下的、三十年前的回波中读出了这座钟。 The planet you cannot see, and cannot stand on, turns out to have been telling you how recently it moved, in the one language it had left, which is the shape of its own scars.这颗你看不见、也无法站立其上的行星,原来一直在用它仅剩的一种语言告诉你它在多近的过去曾经活动过,而那种语言,就是它自己伤痕的形状。

Check your understanding

Try answering before revealing — these are the points the episode turns on.

1. Crater counting suggested Venus's surface is young. Why does that already argue against calling it a 'dead' planet, even before this study?

A geologically dead world just sits there collecting impact craters over billions of years, the way the Moon has. Venus's surface has far too few craters for its age, so most of it must have been repaved within the last few hundred million years. That means something internal has been resurfacing it recently. The open question was never whether Venus had been active, but whether it still is today or did its repaving long ago and then went quiet.

2. On Earth a raised ridge is worn down by rain, rivers and wind. Venus has none of those. So why does the study still expect an old rift flank to be flattened?

Because on Venus the flattening comes from inside the rock, not from weather outside it. The crust is so hot that over long times it yields and flows slowly, more like stiff honey than rigid stone. A raised flank sitting on that soft rock cannot support its own weight, so it sags and spreads back down. That is why erosion is beside the point: even with no weather at all, a flank relaxes and broadens once the rifting that raised it stops.

3. How does the shape of a rift flank act as a clock, and what shape means 'young'?

While a rift is actively pulling apart, its shoulders are held up tall, steep and narrow. Once the pulling stops, the flank relaxes, slumps and broadens relatively quickly in geological terms. So steepness maps to age: a high, sharp-edged flank must be young, while a low, wide, gentle one is old. Reading the steepness of the flanks in old radar therefore lets you estimate how recently that stretch of crust was moving, without ever dating a rock directly.

4. Headlines reported Venus rifting at three to ten centimetres per year. Why is it wrong to treat that as a measurement of Venus today?

Nobody observed Venus moving. That rate is an output of the computer model: it is the widening speed that would produce flanks shaped like the ones seen in the old radar. It is an inference about what speed fits the shapes, not a clocked motion. Treating it as a live figure confuses 'the model implies this rate at some point recently' with 'Venus is measured to be moving this fast now,' which is a much stronger claim than the data support.

5. The scientists say 'recent,' the headlines say 'right now.' What is the actual gap between those, and what would close it?

In this field 'recent' can mean anytime within the last tens of millions of years, so it points to the recent past, not necessarily the present instant. The steep flanks show the rifting is young; they do not show it is happening at this moment. Closing the gap needs a real measurement of motion over time, which is why a future mission like ESA's EnVision matters: fresh radar taken years apart could reveal whether these same flanks are actually shifting, turning a strong inference into an observation.

6. The whole result came from re-examining thirty-year-old data from a spacecraft dead since 1994. What broader lesson does that carry about how discoveries happen?

It shows that a new idea can extract new facts from old observations without any new instrument. The Magellan radar had been available for decades, but no one had a model that turned flank steepness into an age. Once that model existed, the same echoes yielded a conclusion nobody had drawn. Re-reading existing data with a sharper question is often as productive as collecting more, and far cheaper than flying a new mission.

Further reading

  1. Recent active rifting on Venus revealed by wide rift flank uplifts (Nature Geoscience)The primary paper, by Xi Yang, Taras Gerya and colleagues. Abstract free; full text paywalled.
  2. Young rift flanks suggest Venus remains tectonically active (Phys.org)Free. Clear summary of the mechanism and the numbers, with the authors' own caution about timing.
  3. New findings about Venus confirm the planet's tectonic activity (University of Freiburg)Free press release. Co-author Anna Gülcher explains the flattening-by-relaxation idea and notes ESA's EnVision mission will test it.
  4. Venus Rift Valleys Growing at Centimeters Per Year, 3D Simulations Show (Tech Times)Free. A good example of the 'growing centimetres per year' framing the episode argues overshoots the evidence.
  5. Venus Isn't (Geologically) Dead (Scientific American)Background on why the 'dead twin' picture was always shaky, and on coronae. Partly paywalled.
  6. Study Finds Venus' 'Squishy' Outer Shell May Be Resurfacing the Planet (NASA JPL)Free. Explains the young-surface / crater-counting logic and the resurfacing debate that predates this paper.
Episode 005

The hat, and the paperwork

Gunther von Hagens is remembered as a showman who cheapened death; in fact he revived what public anatomy was for four centuries, and the thing to hold against him is not the spectacle you can see but the consent you cannot

2026-08-02

Gunther von Hagens, inventor of plastination and creator of the Body Worlds exhibitions, died in late July 2026 at eighty-one. This episode argues that the popular verdict on him is doubly wrong. The black fedora he always wore was a quotation of Rembrandt's Anatomy Lesson of Dr Tulp, and his public dissections were not a degradation of anatomy but a return to what anatomy was from Vesalius onward: a civic spectacle. The reaction that matters is not disgust at the display but the invisible question of where the bodies came from, and on that the record is mixed and must be stated honestly, distinguishing what he admitted from what he beat in court. The twist is that the man remembered as a ghoul built a consent-based donor program of nearly twenty thousand people, while the worst provenance scandal belongs to his imitators.

🎙 Listen · 9:28
Transcript

Follows the audio as it plays — tap any sentence to jump there.

There is a photograph you have probably seen, even if you never learned the name attached to it.有一张照片你很可能见过,即便你从未知道与它相关的那个名字。 A man in a black fedora, standing beside a human body that has been stripped of its skin and posed as though caught mid-stride, every muscle laid open like the pages of a book.一个戴着黑色礼帽的男人,站在一具被剥去皮肤的人体旁,这具人体被摆成仿佛正迈步行走的姿态,每一块肌肉都像书页一样摊开着。 The man is Gunther von Hagens. He died in the last days of July, at eighty-one, after years of Parkinson's disease.这个男人就是冈瑟·冯·哈根斯(Gunther von Hagens)。他在七月的最后几天去世,享年81岁,此前多年患有帕金森病。 And I want to spend these ten minutes arguing that almost everything you have been told about him is either wrong or aimed at the wrong target.接下来的这十分钟,我想论证的是:关于他,你所听到的几乎一切,要么是错的,要么是指向了错误的对象。

Start with the hat, because the hat is the whole argument in miniature. Most people read it as showmanship. A ghoul in a costume.先从那顶帽子说起,因为这顶帽子正是整个论点的缩影。大多数人把它看作作秀,一个身着戏服的食尸鬼。 It is in fact a quotation.但它其实是一种引用。 It is the hat worn by the surgeon in Rembrandt's painting The Anatomy Lesson of Doctor Tulp, in which a group of seventeenth century gentlemen lean in to watch a corpse being dissected.那是伦勃朗(Rembrandt)画作《杜尔博士的解剖课》(The Anatomy Lesson of Doctor Tulp)中那位外科医生所戴的帽子——画中一群17世纪的绅士俯身观看一具尸体被解剖。 Von Hagens wore that hat every time he appeared in public for thirty years.三十年间,冯·哈根斯每次公开露面都戴着那顶帽子。 He was telling you, if you cared to read it, exactly what he thought he was doing. Not inventing a freak show. Rejoining a tradition.只要你愿意去解读,他是在明确告诉你他认为自己在做什么:不是发明一场怪物秀,而是重新接续一个传统。

But let me give you the man first, because the life is stranger than the exhibitions.不过还是先说说这个人本身,因为他的人生比那些展览更离奇。

He was born in nineteen forty-five, in what was then German-occupied Poland, five days before his family loaded him into a cart and fled west from the advancing Red Army.他出生于1945年,出生地是当时被德国占领的波兰;出生五天后,他的家人就把他装上一辆马车,向西逃离步步逼近的红军。 His birth name was Liebchen.他出生时的姓氏是李布兴(Liebchen)。 He grew up in East Germany and went to medical school in Jena, and in nineteen sixty-nine he tried to escape to the West across the border into Austria.他在东德长大,在耶拿(Jena)读医学院,1969年他试图越过边境逃往西方、进入奥地利。 He was caught. He spent two years in an East German prison for it.他被抓住了,为此在东德的监狱里待了两年。 He got out only because the West German government did what it quietly did in those years. It paid.他能出狱,只是因为西德政府做了那些年里它悄悄在做的事:它出了钱。 Forty-three thousand marks, for one political prisoner.43000马克,换一名政治犯。 He crossed to the other side, finished his training at Heidelberg, and changed his name.他去到了另一边,在海德堡(Heidelberg)完成了训练,并改了名字。

So this is a man who had already been, quite literally, purchased out of confinement. Keep that in mind.所以,这是一个几乎可以说字面意义上被赎出牢笼的人。请记住这一点。 It matters for the end of the story. Now the invention.这对故事的结尾很重要。现在说说那项发明。

In nineteen seventy-seven, working in a pathology lab, von Hagens was watching how anatomical specimens were embedded in plastic.1977年,在一间病理实验室工作时,冯·哈根斯观察着解剖标本是如何被包埋进塑料里的。 The plastic went around the tissue, sealing it in a block. And he had the thought that turned out to be his life.塑料包裹在组织外面,把它封在一个块体里。于是他冒出了一个后来成为他毕生事业的念头。 What if the plastic went inside the tissue instead. What if you could push it into every cell, so that the body itself became the plastic.如果让塑料进入组织内部呢?如果能把它压进每一个细胞,让身体本身变成塑料呢? On his thirty-second birthday, that January, he plastinated a human kidney. He filed the patents over the next three years.那年一月,在他32岁生日那天,他对一枚人体肾脏进行了塑化。接下来的三年里,他陆续申请了相关专利。

Here is what the process actually is, because you should be able to picture it. Your body is mostly water.下面说说这个过程究竟是怎样的,因为你应该能够在脑中想象出来。你的身体大部分是水。 Roughly sixty to seventy percent of you, by weight, is water, and a good deal of the rest is fat. Both of those rot.按重量算,你身体的大约60%到70%是水,剩下的相当一部分是脂肪。这两者都会腐烂。 So the first thing you do is stop the decay and dissect the body to expose whatever you want to show.所以你要做的第一件事,是终止腐败过程,并解剖遗体,把你想展示的部分暴露出来。 Then you drop it into a bath of cold acetone, around twenty-five degrees below zero.然后把它浸入一浴冷丙酮(acetone)中,温度大约在零下25度。 The acetone slowly trades places with the water, molecule for molecule, until the tissue is soaked in acetone instead.丙酮一个分子一个分子地慢慢与水置换,直到组织里浸透的变成丙酮。 Warm the acetone up and it dissolves the fat as well.把丙酮加温,它还会把脂肪一并溶解掉。 Now you have a body with no water and no fat in it, only acetone sitting in all the empty spaces. 现在你得到了一具没有水、也没有脂肪的身体,只有丙酮填充在所有空隙里。

Then comes the clever step, the one that gives the technique its name.接下来是巧妙的一步,也正是这一步给这项技术起了名字。 You submerge the specimen in a bath of liquid silicone and put the whole thing in a vacuum chamber.你把标本浸入液态硅胶浴中,再把整个装置放进真空室里。 Under vacuum, the acetone boils away even though it is cold, and as each pocket of acetone turns to vapour and escapes, the vacuum pulls silicone in behind it to take its place.在真空下,丙酮即便是冷的也会沸腾蒸发,而每一处丙酮化为蒸气逸出时,真空就把硅胶吸进去填补它留下的空缺。 The plastic is quite literally sucked into the vacancy left by the body's own fluids.塑料几乎是被字面意义上地吸入了身体自身液体腾出的空位。 When every space is filled, you pose the specimen, pin it, and cure the silicone with a gas until it sets hard.当每一处空隙都被填满后,你为标本摆好姿势、用针固定,再用一种气体使硅胶固化直到变硬。 A single whole body can take fifteen hundred hours.处理单具完整的遗体可能要花 1500 个小时。 What you are left with does not decay, does not smell, and can be handled in a bright room by a child. That is the invention.你最终得到的东西不会腐烂,不会有气味,可以在明亮的房间里让一个孩子拿在手里。这就是这项发明。 It is genuinely elegant, and it is the reason we are talking about him at all. Now, the tradition.它确实很精巧,也正是我们之所以谈起他的原因。现在来说传统。

The common charge against von Hagens is that he dragged the dignity of anatomy down into spectacle, that he made a circus of the dead.针对冯·哈根斯的常见指责是,他把解剖学的尊严拉低成了一场表演,说他把逝者变成了马戏团。 I think that charge has the history exactly backwards. For four hundred years, dissection was spectacle, and was meant to be.我认为这个指责把历史完全弄反了。四百年来,解剖一直就是表演,而且本就应当如此。 Vesalius, the founder of modern anatomy, performed in front of crowds.现代解剖学的奠基人维萨里,就是当众操作的。 Cities built anatomy theatres, steep wooden amphitheatres with the body on a turntable at the bottom and hundreds of citizens ringed above it, and a public dissection was a civic event, a thing you attended in your good coat.城市建起了解剖剧场——陡峭的木制圆形阶梯厅,底部的转台上摆着遗体,上方环绕着数百名市民,一场公开解剖是一桩市民盛事,是你穿上体面外套去参加的活动。 Rembrandt painted one. That is the hat.伦勃朗画过一幅。那就是那顶帽子。 Anatomy became a private, hidden, professional affair only recently, walled off inside medical schools.解剖成为一件私密的、隐蔽的、专业的事务,只是近来才有的事,被封锁在医学院内部。 What von Hagens did was not a break from the tradition. It was a return to it. You can dislike the return.冯·哈根斯所做的并不是对传统的背离。那是一种回归。你可以不喜欢这种回归。 But if you attack him for turning anatomy into a public show, you are attacking him for restoring the thing anatomy originally was.但如果你因为他把解剖变成公开展演而攻击他,你其实是在因为他恢复了解剖最初本来的样子而攻击他。

He pressed that point hard, and sometimes recklessly.他把这一点推得很用力,有时甚至到了鲁莽的地步。 In November of two thousand and two, in London, he performed the first public autopsy in Britain in one hundred and seventy years, in front of five hundred people, with the cameras running.2002 年 11 月,在伦敦,他进行了英国 170 年来的首次公开尸检,五百人在场,摄像机全程开着。 The official Inspector of Anatomy sent him a letter warning that it broke the law. The Metropolitan Police came.官方的解剖监察官给他寄来一封信,警告说这违反了法律。伦敦警察厅来了人。 They stood and watched, and they did not stop it. He wanted the confrontation. He believed the public had a right to see inside itself.他们站着旁观,并没有阻止它。他要的就是这场对峙。他相信公众有权看到自身的内部。

So far this is a defence. Now I have to turn it around, because there is a real charge here, and it is not the one people usually make.到目前为止这算是一种辩护。现在我得反过来讲,因为这里存在一个真实的指控,而且并不是人们通常提出的那一个。

The thing that makes people squeamish about von Hagens is the sight of it.让人们对冯·哈根斯感到不适的,是眼前所见的那一幕。 A real human corpse, skinned, posed holding its own skin, or sliced into leaves.一具真实的人类遗体,被剥去了皮,摆着姿势拿着自己的皮,或者被切成一片片。 That reaction is strong and it is honest, but I want to suggest it is aimed at the wrong thing. The disgust is about the display.那种反应很强烈,也很真诚,但我想指出,它指向了错误的对象。这种厌恶针对的是展示本身。 The question that actually matters is invisible, and produces no disgust at all, and that is the question of where the bodies came from and whether the people inside them agreed to be there.真正重要的那个问题是看不见的,也丝毫不引起厌恶,那就是这些遗体从何而来、以及身在其中的人是否同意被摆在那里的问题。

And on that question the record is genuinely mixed, so let me separate what is established from what is only alleged.而在这个问题上,记录确实是好坏参半的,所以让我把已经确证的与仅仅是被指控的区分开来。 Von Hagens ran a plastination facility in Dalian, in China, next to which sat prison camps.冯·哈根斯在中国大连经营着一家塑化工厂,紧挨着它的是一些劳改营。 He admitted receiving bodies whose origins he could not verify.他承认收到过一些无法核实来源的遗体。 Two bodies he received from a Chinese university had bullet holes in their skulls.他从一所中国大学收到的两具遗体,颅骨上有弹孔。 In two thousand and four he returned seven corpses to China after conceding they might have come from executed prisoners.2004 年,他在承认这些遗体可能来自被处决的囚犯后,向中国退回了七具遗体。 That much he acknowledged.这一点他是承认的。 What was never proven, and what he fought in court and won, was the stronger claim that the bodies actually on display in his exhibitions were executed prisoners.从未被证实、而他在法庭上抗争并最终胜诉的,是那个更强的指控——即在他展览中实际展出的尸体来自被处决的囚犯。 He said the Chinese bodies were used only as teaching models and never exhibited, and a German court barred the magazine Der Spiegel from asserting otherwise.他说那些中国尸体只用作教学模型,从未被展出,而一家德国法院禁止《明镜》周刊做出相反的断言。 So the honest verdict is uncomfortable in both directions. He handled bodies he could not account for.所以诚实的结论在两个方向上都令人不安。他经手了一些他无法说清来源的尸体。 But the specific charge that his shows were built from the executed was not made to stick.但那个具体的指控——即他的展览是用被处决者的尸体做成的——并未被坐实。

And here is the part that almost no one who calls him a ghoul knows.而下面这一点,几乎没有一个骂他是食尸鬼的人知道。 Von Hagens built a formal body donation program, with living people signing consent while alive.冯·哈根斯建立了一套正式的遗体捐赠项目,由活着的人在生前签署同意书。 By the end there were close to twenty thousand registered donors, and some twenty-eight hundred of them had died and entered the collection with their documented agreement.到最后,登记的捐献者接近 2 万人,其中约 2800 人已经去世,并凭其有据可查的同意进入了收藏。 Meanwhile the imitators, the near identical travelling shows with names like Bodies The Exhibition, run by other companies, are the ones that used unclaimed Chinese bodies with no consent at all.与此同时,那些模仿者——由其他公司经营、名字诸如《Bodies The Exhibition》的几乎一模一样的巡回展——才是那些完全未经任何同意就使用无人认领的中国尸体的人。 One of them settled with the New York Attorney General and admitted, in writing, that it could not verify its bodies were not those of executed prisoners.其中一家与纽约州总检察长达成和解,并以书面形式承认,它无法核实其尸体并非来自被处决的囚犯。 So the man remembered as the one who did this to the dead is the one who built the paperwork of consent.所以,这个被人们记作对死者做出此事的人,恰恰是那个建立了同意书文件制度的人。 The worst of the trade belongs to the copies. He knew he was dying for years.这门交易中最糟糕的部分属于那些仿冒者。他多年来一直知道自己将不久于人世。

Parkinson's, the same disease that had taken his motor control while his mind stayed sharp.帕金森病——正是这种病夺走了他对身体运动的控制,而他的头脑却依然清醒。 And he asked that his own body be plastinated by his wife, Angelina Whalley, who has run the institute, and posed at the entrance of the exhibition, hand outstretched, hat on, to greet the people coming in.他要求由他的妻子安吉丽娜·瓦利——一直经营着该研究所的人——将他自己的遗体进行塑化,并摆放在展览入口处,伸出手、戴着帽子,迎接进来的人们。 A man who was once bought out of a prison for forty-three thousand marks, ending as the one specimen in the room who chose, in full knowledge, to be there.一个曾以 4.3 万马克被从监狱赎出的人,最终成为展厅里唯一一具在完全知情的情况下自愿留在那里的标本。

So here is the corrected sentence. Von Hagens was not a showman who cheapened death.所以,修正后的说法是这样的。冯·哈根斯不是一个把死亡变得廉价的表演者。 He was an anatomist who reopened it to the public, the way it had been for centuries, and the thing to hold against him is not the spectacle you can see.他是一位解剖学家,像几个世纪以来那样,把死亡重新向公众开放,而应该拿来指责他的,不是你能看到的那种景观。 It is the paperwork you cannot.而是你看不到的那些文件。

Check your understanding

Try answering before revealing — these are the points the episode turns on.

1. Why is it a mistake to attack von Hagens for turning anatomy into a public spectacle?

Because for roughly four centuries anatomy was a public spectacle by design. Vesalius dissected in front of crowds, cities built anatomy theatres with the body on a turntable ringed by hundreds of citizens, and Rembrandt painted one such lesson. Anatomy became a private, walled-off professional activity only recently. So the charge that he dragged anatomy down into spectacle has the history backwards: he was returning it to its original public form, not inventing a debasement. You can dislike the return, but the criticism has to be aimed correctly.

2. In the plastination process, what actually drives the silicone into the tissue, and why must the acetone step come first?

A vacuum drives it. The body is first soaked in cold acetone, which trades places with the water and then dissolves the fat, leaving acetone sitting in every space the fluids used to occupy. The specimen is then submerged in liquid silicone under vacuum. Under low pressure the acetone boils off even while cold, and as each pocket of vapour escapes the vacuum pulls silicone in to fill the vacancy. The acetone is essential because it is a volatile stand-in: water will not boil out under those conditions, but acetone will, and that boiling-and-replacement is the mechanism.

3. The episode says the thing that disgusts people about Body Worlds is not the thing that should worry them. What is the distinction?

The disgust is a reaction to the display: a real corpse, skinned and posed. That reaction is honest but it is aimed at appearances. The question that actually carries the moral weight is invisible and produces no disgust at all: where each body came from and whether the person inside it consented. A tastefully hidden body obtained without consent is worse than a shocking one freely donated, yet our instinct rates them the other way. The argument is that our squeamishness points at the wrong variable.

4. What did von Hagens admit about the Chinese bodies, and what did he manage to deny in court? Why does the difference matter?

He admitted running a facility in Dalian, receiving bodies whose origins he could not verify, receiving two with bullet holes in the skull, and returning seven corpses in 2004 because they might have been executed prisoners. What he denied, and won on, was the stronger claim that the bodies displayed in his exhibitions were executed prisoners; a German court barred Der Spiegel from asserting it. The difference matters because honesty requires separating what a person conceded from what was only alleged. The record convicts him of careless custody, not of the specific atrocity he is often accused of.

5. Why is it ironic that von Hagens is the figure remembered as a ghoul?

Because he is the one who built the consent machinery. His Body Worlds bodies come from a formal donor program in which living people signed up, close to twenty thousand of them, with some twenty-eight hundred having died and entered the collection by documented agreement. The near-identical imitator shows, run by other companies, are the ones that used unclaimed Chinese bodies without consent, one settling with the New York Attorney General and admitting it could not verify the bodies were not those of executed prisoners. Public memory pinned the crime on the innovator and largely spared the copies.

6. What is the significance of von Hagens asking to be plastinated and posed at the entrance of his own exhibition?

It closes the argument about consent by making him the one specimen who chose, in full knowledge, to be there. A man who was once literally bought out of an East German prison for forty-three thousand marks ends as a body that entered the collection by his own explicit will, greeting visitors with an outstretched hand and the Rembrandt hat. It is also a statement of what he believed: that there is no indignity in a body being seen, only in a body being taken.

Further reading

  1. Gunther von Hagens — WikipediaFree. The fullest sourced timeline: the East German prison and ransom, the patents, the Rembrandt hat, the 2002 London autopsy, and the China and Kyrgyzstan provenance cases with their legal outcomes.
  2. The Plastination Process — von Hagens PlastinationFree, and it is his own institute, so read it as the maker's account. Clear on the four steps: fixation, cold acetone dehydration, forced impregnation under vacuum, and curing.
  3. Body Worlds — WikipediaFree. Background on the donor program and the exhibitions, and a useful counterweight to the more defensive institute pages.
  4. Bodies: The Exhibition — WikipediaFree. The imitator run by Premier Exhibitions, its use of unclaimed Chinese bodies, and the New York Attorney General settlement in which it admitted it could not verify the bodies were not those of executed prisoners.
  5. Donors sign up to have bodies dissected, displayed — CNNFree. On the living-donor consent program that is easy to overlook when the only image you have of von Hagens is the posed corpse.
  6. Exclusive: Secret Trade in Chinese Bodies — ABC NewsFree. The investigative reporting on the Dalian trade in cadavers; useful for the provenance charges, but read it alongside the Wikipedia entry on which claims were proven and which von Hagens beat in court.
  7. The Anatomy Lesson of Dr Nicolaes Tulp — WikipediaFree. The 1632 Rembrandt painting the hat quotes, and a window onto the era when public dissection was a paid civic event rather than a scandal.
Episode 004

The man who never changed his mind

Geoffrey Hinton is read as a builder who turned against his creation. He isn't. He has held one idea since the 1960s, and the warning is that idea reaching its conclusion

2026-08-01

A profile of Geoffrey Hinton: the Boole family inheritance, twenty years of being wrong in public, the 2012 result that flipped the industry in a year, the Turing Award and a contested Nobel, and the 2023 resignation. The argument is that none of it is a change of heart. He spent fifty years insisting that minds are machines, he won, and winning is exactly what frightens him. Includes the parts profiles usually leave out: the credit he does not claim, the colleagues who disagree with him, and the late bet that failed.

🎙 Listen · 10:12
Transcript

Follows the audio as it plays — tap any sentence to jump there.

On Tuesday, in Las Vegas, a seventy-eight year old man will walk onto a conference stage alongside Fei-Fei Li and Andrew Ng. 周二,在拉斯维加斯,一位七十八岁的老人将与李飞飞和吴恩达一同走上一个会议的舞台。 He is billed, as he always is now, as the godfather of artificial intelligence.如今人们介绍他时,一如既往地称他为人工智能的教父。 And he will almost certainly spend his time on stage explaining why he is worried about it. Most people read that as a change of heart.而他几乎肯定会把台上的时间用来解释自己为何为此感到忧虑。大多数人把这理解为一种态度的转变。

The man who built the thing, turning against the thing. A late conversion, a deathbed confession.那个造出了这东西的人,如今转过身来反对它。一次迟来的皈依,一场临终的忏悔。

I want to argue that it is nothing of the kind.我想说,这根本不是那么回事。 Geoffrey Hinton has believed one thing, more or less continuously, since he was a student in the nineteen sixties.自二十世纪六十年代还是学生起,Geoffrey Hinton 就大体上从未间断地相信着同一件事。 For fifty years that belief made him a crank. Then it made him famous. And then, followed through to its end, it frightened him.五十年里,这个信念让他被当成怪人。后来,它让他成名。再后来,当它被贯彻到底,它让他感到害怕。 The warning is not a reversal. It is the same idea arriving at its destination. Start with the family, because it is genuinely absurd.这番警告不是一次反转。它是同一个想法抵达了它的终点。先从他的家世说起,因为这实在荒诞得离谱。

His great-great-grandfather was George Boole. That Boole. Boolean logic, the true-and-false algebra that every computer on earth runs on.他的高外祖父是 George Boole。就是那个 Boole。布尔逻辑,那套真与假的代数,地球上每一台计算机都靠它运转。

Boole's wife, Mary Everest Boole, was herself a mathematician, and her uncle was George Everest, the surveyor the mountain is named after.Boole 的妻子 Mary Everest Boole 本人也是数学家,而她的叔父是 George Everest,那座山就是以这位测量师命名的。 Which is why Geoffrey Hinton's middle name is Everest.这就是为什么 Geoffrey Hinton 的中间名叫 Everest。 His great-grandfather, Charles Howard Hinton, was a mathematician who wrote about the fourth dimension and gave us the word tesseract.他的曾外祖父 Charles Howard Hinton 是一位数学家,写过关于第四维度的著作,并给了我们 tesseract(超立方体)这个词。 His father was an entomologist and a Fellow of the Royal Society.他的父亲是一位昆虫学家,也是英国皇家学会的会士。

Hinton has said that growing up in that family, the message was fairly clear.Hinton 说过,在那样一个家庭里长大,传递出的信息相当明确。 You could become an academic, or you could be a disappointment. At Cambridge he could not settle.你要么成为一名学者,要么就是个让人失望的人。在剑桥,他一直安顿不下来。

He started in physiology and physics, switched to philosophy, switched again, and finally took his degree in experimental psychology.他从生理学和物理学起步,转去哲学,又转了一次,最后拿的是实验心理学的学位。 What he was actually chasing, in all that switching, was one question: how does the brain work. Physics did not answer it.在所有这些转向背后,他真正追逐的是一个问题:大脑是如何运作的。物理学没有回答它。 Philosophy did not answer it. Psychology, he decided, did not really answer it either.哲学没有回答它。他认定,心理学其实也没有真正回答它。

So he went to Edinburgh, to do a doctorate in artificial intelligence, which he finished in nineteen seventy-eight.于是他去了爱丁堡,攻读人工智能的博士学位,并于一九七八年完成。 His supervisor thought neural networks were a dead end. So did the field.他的导师认为神经网络是一条死路。整个领域也这么认为。 Nine years earlier, Marvin Minsky and Seymour Papert had published a book that laid out what simple neural networks could not do, and the effect was to more or less end the subject.九年前,Marvin Minsky 和 Seymour Papert 出版了一本书,阐明了简单神经网络做不到哪些事,其效果差不多终结了这个课题。 Funding went elsewhere. Serious people worked on something else. That something else was symbolic AI.经费流向了别处。严肃的人去做别的东西。那别的东西就是符号主义 AI。

The idea was that intelligence is reasoning, reasoning is manipulating symbols according to rules, so if you want a machine to be intelligent you write down the rules and the facts.其思路是:智能就是推理,推理就是按照规则操纵符号,所以如果你想让一台机器有智能,你就把规则和事实写下来。 Encode what a doctor knows and you get a machine that diagnoses. Hinton thought this was backwards, and his reason was biological.把一位医生所知道的东西编码进去,你就得到一台会诊断的机器。Hinton 认为这是本末倒置,他的理由是生物学上的。

There are no rules written anywhere in your head.你脑袋里任何地方都没有写着规则。 There is a very large number of cells, connected to each other, and the connections get stronger or weaker with experience.有的是数量极其庞大的细胞,彼此相连,而这些连接会随着经验变强或变弱。 That is all there is. If that lump of tissue can recognize your grandmother, then recognizing your grandmother does not require rules. 如此而已。如果那一团组织能认出你的祖母,那么认出你的祖母就不需要规则。 It requires the right connection strengths. And nobody could write those down by hand, so the machine would have to learn them.它需要的是恰当的连接强度。而没有人能靠手工把这些写下来,所以机器只能自己去学。

That is the belief. Thinking is a physical process.这就是那个信念。思考是一个物理过程。 Whatever the brain is doing, it is doing it with adjustable connections, and therefore a machine with adjustable connections could do it too.无论大脑在做什么,它都是用可调节的连接在做,因此一台拥有可调节连接的机器也能做到。

In nineteen eighty-six he published a paper with David Rumelhart and Ronald Williams, in Nature, on backpropagation.1986 年,他与 David Rumelhart 和 Ronald Williams 在《自然》上发表了一篇关于反向传播(backpropagation)的论文。 This is the one everyone cites. And here I want to be careful, because the popular version gets it wrong.这是所有人都会引用的那一篇。这里我要谨慎一点,因为流行的说法把它搞错了。 Hinton did not invent backpropagation.Hinton 并没有发明反向传播。 The mathematics had been worked out before, by Seppo Linnainmaa in nineteen seventy, and applied to this kind of problem by Paul Werbos in the seventies.其数学在此之前就已被推导出来——1970 年由 Seppo Linnainmaa 完成,并在七十年代由 Paul Werbos 应用到这类问题上。 What the nineteen eighty-six paper did was show that if you train a network this way, the middle layers learn useful internal representations on their own.1986 年那篇论文所做的,是证明如果用这种方式训练一个网络,中间层会自行学到有用的内部表示。 Nobody designs them. They emerge.没有人去设计它们。它们是自发涌现的。 That was the demonstration that mattered, and Hinton has been consistently honest about the earlier credit, which is more than the field has always been.这才是真正重要的证明,而 Hinton 一直诚实地承认前人的功劳,这一点比这个领域一贯的做法要好。

Then came twenty years of being wrong in public. Through the nineteen nineties and into the two thousands, neural networks lost.接下来是二十年当众被判定为错误的日子。整个九十年代直到两千年代,神经网络输了。

Support vector machines worked better on the problems people cared about, and they came with mathematical guarantees.在人们关心的问题上,支持向量机(SVM)表现更好,而且还带有数学上的保证。 Neural networks were slow, needed data nobody had, and were considered slightly embarrassing.神经网络又慢,需要谁都没有的数据,还被认为有点上不了台面。 Papers were rejected because of the words in the title. Students were advised to work on something with a future. Hinton moved to Canada.论文因为标题里的字眼就被拒。学生被劝去做些有前途的方向。Hinton 搬去了加拿大。

Partly because of where American AI money came from, which was largely military and which he did not want.一部分原因是美国的 AI 经费来自何处——大多来自军方,而这是他不想要的。 Partly because a Canadian institute was willing to fund a small group doing unfashionable work for a long time with no obvious payoff.一部分原因是一家加拿大机构愿意长期资助一个小团队去做不时髦、没有明显回报的工作。 He, Yann LeCun and Yoshua Bengio kept going. It was not a large community. The payoff came in twenty twelve.他、Yann LeCun 和 Yoshua Bengio 坚持了下来。这个圈子并不大。回报在 2012 年到来。

Two of his graduate students, Alex Krizhevsky and Ilya Sutskever, entered a network in the ImageNet competition, a contest for classifying photographs.他的两名研究生 Alex Krizhevsky 和 Ilya Sutskever 把一个网络送进了 ImageNet 竞赛——一项给照片分类的比赛。 It did not edge out the competition. It demolished it, cutting the error rate by a margin that made the result look like a mistake.它不是险胜对手,而是把对手碾碎了,将错误率降低的幅度之大,让这个结果看起来像是出了差错。 The thing that changed was not really the idea. The idea was thirty years old.真正改变的并不是这个想法。这个想法已经有三十年历史了。 What changed was that there were now enough images to learn from, and graphics cards fast enough to do the arithmetic.改变的是,如今有了足够多可供学习的图像,也有了快到足以完成这些运算的显卡。

A few months later, Google paid roughly forty-four million dollars for a company with three employees, no product and no revenue.几个月后,Google 为一家只有三名员工、没有产品也没有营收的公司支付了大约 4400 万美元。 The company was Hinton and those two students. Within about a year, every large technology firm on earth had reorganized around this.这家公司就是 Hinton 和那两名学生。大约一年之内,地球上每一家大型科技公司都围绕这件事进行了重组。

Then the honours. The Turing Award in twenty eighteen, shared with Bengio and LeCun.然后是各种荣誉。2018 年的图灵奖,与 Bengio 和 LeCun 共享。 And in twenty twenty-four, the Nobel Prize in Physics, shared with John Hopfield, which caused a certain amount of grumbling about whether any of this is physics.以及 2024 年的诺贝尔物理学奖,与 John Hopfield 共享,这引来了一些抱怨,质疑这一切究竟算不算物理学。 Hinton himself seemed to find it funny. And in May of twenty twenty-three, he left Google.Hinton 自己似乎觉得这挺好笑。而在 2023 年 5 月,他离开了 Google。

He was specific about why, and the precision matters. He did not say Google had behaved badly.他明确说明了原因,而这种精确很重要。他并没有说 Google 做得不好。 He said he wanted to be able to talk about the dangers without having to think about how it reflected on his employer.他说他希望能够谈论那些危险,而不必去考虑这会如何影响到他的雇主。 He resigned in order to be free to say something, not to protest something. So what is the something? Here is where the through-line closes.他辞职是为了能自由地说出某些话,而不是为了抗议某件事。那么这某些话是什么?主线就在这里收拢。

If you spend fifty years arguing that minds are machines, you do not get to be surprised when machines start to look like minds.如果你花五十年主张心智就是机器,那么当机器开始看起来像心智时,你就没有资格感到惊讶。

That was the whole thesis. There is no magic ingredient in biological tissue.这就是整个论点。生物组织里并没有什么神奇的成分。 Intelligence is what a sufficiently large learning system does. Hinton won that argument.智能就是一个足够大的学习系统所做的事。Hinton 赢下了这场争论。 What he now says, roughly, is that having won it, he sees no principled reason why the machines should stop at our level, and no good plan for what happens when they do not.他现在大致的说法是,既然赢了,他看不出有什么原则性的理由能让机器止步于我们的水平,也看不到当它们不止步时该怎么办的好方案。

His recent claims are more concrete than the headlines suggest.他近期的主张比新闻标题所暗示的更为具体。 He argues that the length of task a system can complete has been doubling every seven months or so, and asks what that curve looks like in a few years.他认为,一个系统能够完成的任务时长大约每七个月就翻一番,并追问几年之后这条曲线会是什么样子。 On jobs, he has said something notably unusual for a technologist: that mass unemployment plus enormous profits is not a fact about AI but a fact about how we have arranged the economy.在就业问题上,他说了一句对技术专家而言相当反常的话:大规模失业加上巨额利润,并不是关于 AI 的事实,而是关于我们如何安排经济的事实。 It will make a few people much richer and most people poorer. That is a claim about capitalism, not about neural networks.它会让少数人更富,让大多数人更穷。这是一个关于资本主义的论断,而不是关于神经网络的论断。

Now, three honest complications, because a profile that only admires is a press release. First, the title. Godfather of AI, father of AI.接下来是三点如实的复杂之处,因为一篇只有赞美的人物剖析不过是一份新闻稿。第一,头衔。AI 教父、AI 之父。

These are media inventions.这些都是媒体的发明。 Artificial intelligence as a field is older than Hinton's career and much wider than neural networks, and there are researchers, Jürgen Schmidhuber most vocally, who argue that the credit for these ideas has been systematically misassigned.作为一个领域,人工智能比 Hinton 的职业生涯更古老,也远比神经网络宽广,而且有些研究者——其中以 Jürgen Schmidhuber 呼声最高——认为这些思想的功劳被系统性地错误归属了。 The label flattens a crowded history into one face. Second, the experts do not agree.这个标签把一段拥挤的历史压平成了一张面孔。第二,专家们意见并不一致。

Yann LeCun, who shares the Turing Award with him, for the same work, thinks the existential worry is misguided.Yann LeCun 因同样的工作与他共享图灵奖,他认为那种关乎存亡的担忧是被误导的。 When you hear that the people who built this are frightened, remember that some of the people who built this are not.当你听说建造出这一切的人感到恐惧时,请记住,建造出这一切的人中也有一些并不恐惧。

Third, and I think most usefully, Hinton has been wrong recently.第三,而且我认为这一点最有用:Hinton 近来犯过错。 His major late bet was capsule networks, an attempt to fix what he saw as a deep flaw in the systems he helped create. It did not work out.他近期押下的一个重大赌注是 capsule networks(胶囊网络),试图修复他所认为的、他参与创造的那些系统中的一个深层缺陷。结果并不成功。 The field went a different way, with transformers.整个领域走上了另一条路,用的是 transformers。 Which is worth holding onto, because the stubbornness that kept him working on neural networks through twenty years of ridicule is the same stubbornness that kept him on capsules.这一点值得记住,因为让他在二十年的嘲笑中坚持研究神经网络的那份固执,正是让他坚持研究胶囊网络的同一份固执。 Being right for decades when everyone disagrees does not come from a faculty that only fires when you are right.在所有人都不认同时,能连续几十年正确,这并非来自某种只在你正确时才启动的能力。

One more thing, briefly, because it is part of the life and it is usually left out.还有一件事,简短提一下,因为它是这段人生的一部分,却通常被略去。 His second wife, Rosalind, died of ovarian cancer in nineteen ninety-four.他的第二任妻子 Rosalind 于 1994 年死于卵巢癌。 His third wife, Jacqueline, died of pancreatic cancer in twenty eighteen. He raised two children through the first of those.他的第三任妻子 Jacqueline 于 2018 年死于胰腺癌。在前一场变故中,他独自把两个孩子抚养长大。 The wilderness decades were not only professional. What I keep coming back to is that this is not the story of a man who changed his mind.那些困顿的岁月不只体现在职业上。我一再想到的是,这并不是一个改变了想法的人的故事。

It is the story of a man who did not, for fifty years, against sustained evidence that he should. He argued the mind is a machine.这是一个五十年里都没有改变想法的人的故事——尽管持续有证据表明他应该改变。他主张心智是一台机器。 He was right. And being right is precisely what worries him. Next time, something else entirely. Thanks for listening.他是对的。而正是这份正确让他担忧。下一期,聊些完全不同的东西。感谢收听。

Check your understanding

Try answering before revealing — these are the points the episode turns on.

1. The episode argues Hinton's warnings are not a change of heart. What is the single belief that connects the 1970s heretic to the 2023 resignation, and why does it lead where it leads?

The belief is that thinking is a physical process: there are no rules written anywhere in a brain, only a very large number of connections whose strengths change with experience, and therefore intelligence is something a learning system does rather than something a designer programs. That view was heresy when symbolic AI dominated. But if you hold it and you are right, two consequences follow without any further assumption. There is no special ingredient in biological tissue that machines lack. And there is no principled ceiling at human level, because nothing in the argument says the quantity of connections or the quality of learning stops there. The warning is that conclusion stated out loud, not a new position.

2. Hinton did not invent backpropagation. What did the 1986 paper actually establish, and why does the distinction matter for how you read his reputation?

The mathematics predates him: Linnainmaa worked it out around 1970 and Werbos applied it to this class of problem in the 1970s. What the 1986 paper with Rumelhart and Williams showed was that training a multi-layer network this way causes the hidden layers to develop useful internal representations on their own, without anyone designing them. That is a claim about what learning produces, not about an algorithm. The distinction matters because the popular story compresses a long lineage into one inventor, and because Hinton himself has been consistent about the earlier credit. A reputation that survives being described accurately is worth more than one that needs the myth.

3. The idea was thirty years old in 2012. So what actually changed to make AlexNet possible, and what does that suggest about how progress in this field works?

Two things that were not ideas: enough labelled images to learn from, in ImageNet, and graphics processors fast enough to make the arithmetic tractable. The architecture and the training method were not new in kind. This suggests a pattern worth carrying into any judgement about the field: capability jumps here have often come from scale and hardware crossing a threshold rather than from conceptual breakthroughs, which means an approach can look dead for decades while being merely early. It also cuts the other way, as a caution against assuming that today's limitations are conceptual when they might just be waiting for the next threshold.

4. The episode insists on mentioning that Yann LeCun disagrees with Hinton about existential risk. Why is that a load-bearing detail rather than a courtesy?

Because the persuasive force of Hinton's warning is often taken to be that the people who built the technology are afraid of it, which implies a consensus among those best placed to know. LeCun shares the Turing Award with Hinton for the same body of work and does not share the worry. So the premise is false as stated. The honest version is narrower: some of the field's founders are alarmed and others are not, which means the disagreement has to be settled on arguments rather than on authority. Anyone citing Hinton's credentials as the reason to believe him should be equally moved by LeCun's.

5. Why does the episode bring up capsule networks, which failed, in a profile of someone's successes?

Because it identifies the trait rather than flattering the man. Hinton kept working on neural networks through roughly twenty years in which the field considered them a dead end, and he was vindicated. He then bet heavily on capsule networks as a fix for what he saw as a deep flaw, and the field went to transformers instead. The same persistence produced both. That matters for how you weigh his current warnings: the disposition that made him right for decades is not a faculty that only activates when he happens to be correct, so his track record is evidence about his seriousness rather than a guarantee about this particular claim.

6. Hinton's claim about jobs is that AI will make a few people much richer and most people poorer. What kind of claim is that, and why is it worth separating from his other predictions?

It is a claim about political economy, not about machine learning. The technical prediction is that systems will become capable of more tasks; the further step, that this produces concentrated gains and widespread loss, depends entirely on how ownership, taxation and bargaining power are arranged, not on anything inside the models. Separating the two matters because they are checked differently and can be acted on differently: the first is a forecast about capability that time will settle, while the second is a statement that the harm is a policy choice rather than an inevitability. Conflating them lets a technical authority carry weight on a question where his expertise does not obviously apply.

Further reading

  1. Geoffrey Hinton (Wikipedia)The reliable spine of dates, positions and family. Free.
  2. Rumelhart, Hinton & Williams (1986) — Learning representations by back-propagating errors, Nature 323, 533–536The paper everyone cites. Note what it actually claims: not that backpropagation is new, but that hidden layers learn useful representations by themselves. Paywalled.
  3. The Nobel Prize in Physics 2024 — Hopfield and HintonThe citation, and the committee's reasoning for why they considered this physics. Free.
  4. ACM A.M. Turing Award 2018 — Bengio, Hinton, LeCunThe award all three shared, for the same body of work they now disagree about the consequences of. Free.
  5. Geoffrey Hinton (Encyclopaedia Britannica)A compact, edited biography if you want one source rather than several. Free.
  6. 'Godfather of AI' Geoffrey Hinton predicts 2026 will see AI 'replace many other jobs' (Fortune)His current position in his own words, including the seven-month doubling claim and the argument that the problem is the economic system rather than the technology. Free.
  7. Jürgen Schmidhuber — Critique of Honors 2018 (Turing Award for Deep Learning)The other side of the credit question, argued at length and with citations. Read it as advocacy, but the specific priority claims are checkable. Free.
Episode 003

Four medals, and what they were really for

The 2026 Fields Medals: a needle on a table, the arrow of time, two ways of counting, and a tool borrowed from logic

2026-07-30

On 23 July in Philadelphia, Hong Wang, Yu Deng, John Pardon and Jacob Tsimerman were awarded the Fields Medal. This episode explains what each of them actually did, in plain language: why a puzzle about rotating a needle underpins half of harmonic analysis, where the arrow of time comes from if Newton's laws don't have one, why two independent ways of counting the same curves agreeing is a big deal, and how a tool built by logicians ended up solving a fifty-year-old problem in geometry.

🎙 Listen · 10:02
Transcript

Follows the audio as it plays — tap any sentence to jump there.

Last Thursday, in Philadelphia, four people under the age of forty were handed a gold medal with the face of Archimedes on it.上周四,在费城,四位年龄不到四十岁的人获颁一枚刻有阿基米德头像的金牌。 The Fields Medal is given once every four years, to between two and four mathematicians, and it comes with about ten thousand US dollars, which is famously almost nothing.菲尔兹奖每四年颁发一次,授予两到四位数学家,奖金约一万美元——这笔钱少得出名,几乎不值一提。 The prestige is the whole prize. People call it the Nobel of mathematics, and that is not quite right. The Nobel is for a lifetime of work.荣誉本身才是全部的奖赏。人们称它为数学界的诺贝尔奖,但这并不太准确。诺贝尔奖表彰的是毕生的工作。

The Fields Medal has an age limit: you must not have turned forty. It is explicitly a bet on the future as much as a reward for the past.菲尔兹奖有年龄限制:你不能已满四十岁。它明确地既是对过去的奖励,也是对未来的押注。 Which means the list of winners is also a snapshot of where mathematics thinks it is going.这意味着获奖名单也是数学界对自身走向的一个快照。

This year's four are Hong Wang, Yu Deng, John Pardon and Jacob Tsimerman.今年的四位得主是王虹、邓煜、John Pardon 和 Jacob Tsimerman。 And a few things about the list are worth noticing before we get to the mathematics.在进入数学内容之前,这份名单有几点值得注意。

Hong Wang is the third woman to win in the medal's ninety year history. The first was Maryam Mirzakhani in twenty fourteen.王虹是这枚奖章九十年历史上第三位获奖的女性。第一位是 2014 年的 Maryam Mirzakhani。 The second was Maryna Viazovska in twenty twenty-two. So three, out of about sixty five.第二位是 2022 年的 Maryna Viazovska。所以在约六十五位得主中,只有三位女性。

Wang and Deng are also only the second and third medalists born in mainland China. The first was Shing-Tung Yau, in nineteen eighty-two.王虹和邓煜也只是出生于中国大陆的第二位和第三位得主。第一位是 1982 年的丘成桐。 That is a gap of forty four years. Now, the mathematics.这中间相隔了四十四年。现在,来谈数学。

I want to do these one at a time, because each is a genuinely good story, and two of them you can picture in your head.我想一个一个来讲,因为每一个都是真正精彩的故事,而其中两个你可以在脑海里想象出来。

Start with Hong Wang, because hers is the most visual problem in all of modern mathematics.先从王虹讲起,因为她的问题是全部现代数学中最具画面感的一个。

In nineteen seventeen, a Japanese mathematician named Soichi Kakeya asked a question that sounds like a puzzle from a magazine.1917 年,一位名叫挂谷宗一(Soichi Kakeya)的日本数学家提出了一个听起来像是杂志谜题的问题。 You have a needle of length one, lying on a table.你有一根长度为 1 的针,平放在桌面上。 You want to rotate it a full one hundred and eighty degrees, so it ends up pointing the opposite way.你想把它整整旋转 180 度,让它最终指向相反的方向。 You are allowed to slide it around as you turn. What is the smallest area of table you need to sweep out?在旋转过程中你可以任意滑动它。你需要扫过的桌面面积最小是多少?

The obvious answer is to spin it about its centre, which sweeps a disc. You can do better with a shape like a three pointed star.显而易见的答案是绕它的中心旋转,这样会扫出一个圆盘。用一个类似三角星的形状可以做得更好。 But then a Russian mathematician, Abram Besicovitch, proved something genuinely shocking. There is no smallest area.但随后,俄罗斯数学家 Abram Besicovitch 证明了一件真正令人震惊的事:根本不存在最小面积。 You can turn the needle using as little area as you like. As close to zero as you want. That should bother you.你可以用任意小的面积来旋转这根针,想多接近零就多接近零。这应该让你感到困惑。

The needle has to point in every direction at some moment. Every direction. And yet the total area swept can be essentially nothing.这根针必须在某个时刻指向每一个方向。每一个方向。然而扫过的总面积却可以基本为零。

So area is the wrong way to measure these sets. Mathematicians switched to a subtler ruler called dimension.所以用面积来衡量这些集合是错误的方式。数学家转而使用一把更微妙的尺子,叫做维数。 Not the dimension you learned in school, where a line is one and a plane is two, but a version that allows fractions.不是你在学校学的那种维数——直线是一维、平面是二维——而是一个允许分数的版本。 A set can have zero area and still be, in this finer sense, two dimensional. Or one point five dimensional.一个集合可以面积为零,但在这种更精细的意义上仍然是二维的。或者是 1.5 维的。 The dimension measures how thoroughly the set fills space, even when its area is zero.维数衡量的是这个集合填充空间的彻底程度,即使它的面积为零。

And the conjecture is this: a set that contains a needle pointing in every direction must have full dimension. In the plane, two.而猜想是这样的:一个包含指向每个方向的针的集合必定具有满维数。在平面上,就是二维。 In three dimensional space, three. It can be as thin as you like in terms of area, but it cannot be thin in terms of dimension.在三维空间里,就是三维。它在面积上可以任意地薄,但在维数上不可能薄。

The plane case was settled in nineteen seventy-one. Three dimensions stayed open for another fifty years.平面的情形在 1971 年得到解决。三维的情形又悬而未决了五十年。 In February of twenty twenty-five, Hong Wang and Joshua Zahl proved it.2025 年 2 月,王虹和 Joshua Zahl 证明了它。

Now, why does anyone outside this corner of mathematics care about a needle? Because the Kakeya problem turned out to be a bottleneck.那么,为什么这个数学角落之外的人会关心一根针呢?因为挂谷问题最终被证明是一个瓶颈。 It sits underneath a surprising number of other things. How waves spread out and focus.它支撑着数量惊人的其他问题。波如何扩散和聚焦。 A central question in Fourier analysis about which frequencies you can reconstruct a signal from. Even parts of number theory.傅里叶分析中的一个核心问题——你能从哪些频率重建一个信号。甚至还有数论的一部分。 Terence Tao has described it as something like a Rosetta stone: solve it and you get leverage on a whole family of problems that look unrelated.陶哲轩曾把它描述成某种罗塞塔石碑:解开它,你就能撬动一整族看似毫不相关的问题。 That is why a puzzle about turning a needle on a table is worth a Fields Medal. Second, Yu Deng, and a problem that is even older.这就是为什么一个关于在桌上转动一根针的谜题值得一枚菲尔兹奖。第二位,Yu Deng,以及一个更古老的问题。

In nineteen hundred, in Paris, David Hilbert stood up and listed twenty three problems he thought should occupy the coming century.1900 年,在巴黎,David Hilbert 起身列出了他认为应当占据未来一个世纪的二十三个问题。

The sixth one was different from the others. It was not really a mathematics problem. It asked for physics to be put on a rigorous footing.第六个问题与其他的不同。它其实不是一个数学问题。它要求把物理学建立在严格的基础之上。 In particular: we describe gases in two completely different ways, and nobody had shown the two descriptions were compatible.尤其是:我们用两种完全不同的方式描述气体,而没有人证明过这两种描述是相容的。

Here is the tension, and it is a good one. At the small scale, a gas is just particles bouncing off each other according to Newton's laws.这里有一个张力,而且是个很好的张力。在小尺度上,气体不过是按照牛顿定律相互碰撞的粒子。 Those laws are reversible. Film a collision, run the film backwards, and what you see is still a legal collision.那些定律是可逆的。把一次碰撞拍下来,把影片倒着放,你看到的仍然是一次合法的碰撞。 Nothing in the microscopic physics knows which way time points. But at the large scale, gases obviously do know.微观物理中没有任何东西知道时间指向哪个方向。但在大尺度上,气体显然是知道的。

Heat flows from hot to cold and never the other way. Smoke spreads out and never gathers itself back into the cigarette.热量从热流向冷,从不反过来。烟扩散开去,从不会自己重新聚回到香烟里。 The equations we use at that scale, Boltzmann's equation and the equations of fluid flow, have an arrow of time baked into them.我们在那个尺度上使用的方程,玻尔兹曼方程和流体流动的方程,本身就内建了一支时间之箭。

So where does the arrow come from, if it is not in the underlying rules?那么,如果时间之箭不在底层的规则里,它是从哪里来的?

Boltzmann wrote down the bridging equation in eighteen seventy-two, and was attacked for it, on exactly this point.玻尔兹曼在 1872 年写下了那个桥接方程,并正是在这一点上受到了攻击。 A rigorous derivation had to wait a century.一个严格的推导不得不等了一个世纪。 In nineteen seventy-five, Oscar Lanford finally proved it, and the proof was a landmark, but it had a serious catch.1975 年,Oscar Lanford 终于证明了它,这个证明是一座里程碑,但它有一个严重的缺陷。 It only worked for an extremely short window of time. Roughly, less than the time it takes a typical particle to collide once.它只在极短的一段时间内成立。大致上,短于一个典型粒子碰撞一次所需的时间。 That is not really enough to claim you have explained a gas. Yu Deng, with Zaher Hani and Xiao Ma, removed the time restriction.这其实不足以宣称你解释了气体。Yu Deng,与 Zaher Hani 和 Xiao Ma 一起,去掉了这个时间限制。

Their result holds for as long as the Boltzmann equation itself makes sense.他们的结果在玻尔兹曼方程本身有意义的整个时段内都成立。 And then they pushed further, deriving the equations of fluid mechanics on top of that. A hundred and twenty six years after Hilbert asked.然后他们更进一步,在此之上推导出了流体力学的方程。距 Hilbert 提出问题已过去一百二十六年。

Third, John Pardon. This one is harder to picture, so let me give you the shape of it rather than the details.第三位,John Pardon。这个更难想象,所以让我给你它的轮廓而不是细节。

There is a class of geometric objects called Calabi-Yau threefolds.有一类几何对象叫做 Calabi-Yau 三维流形。 They come up in string theory as candidate shapes for the hidden dimensions of space.它们出现在弦论中,作为空间隐藏维度的候选形状。 A natural thing to ask about such a shape is: how many curves of a given type does it contain?关于这样一个形状,一个很自然的问题是:它包含多少条给定类型的曲线? Think of it as a very sophisticated counting problem. The trouble is that there were two different ways to count.把它想成一个非常精巧的计数问题。麻烦在于有两种不同的计数方式。

One came from symplectic geometry, the other from algebraic geometry. Different definitions, different machinery, different communities.一种来自辛几何,另一种来自代数几何。定义不同,工具不同,群体也不同。 Around two thousand and three, four mathematicians conjectured that the two counts always agree, after a suitable translation between them.大约在 2003 年,四位数学家猜想,经过两者之间合适的转译之后,这两种计数总是一致的。

Pardon proved it. And the reason this matters is more than bookkeeping.Pardon 证明了它。而这件事之所以重要,不只是账目上的对齐。 When two entirely independent methods of counting the same thing always give the same answer, that is strong evidence you are counting something real, rather than measuring an artifact of your own definitions.当两种完全独立的方法在计数同一样东西时总是给出相同的答案,这就是强有力的证据,表明你在计数某种真实的东西,而不是在度量你自己定义的一个人为产物。 It welds two fields together. Pardon, incidentally, has a reputation for this.它把两个领域焊接在了一起。顺带一提,Pardon 在这方面颇有名声。

He proved a famous conjecture about symmetries of three dimensional spaces while he was still an undergraduate.他还在读本科时就证明了一个关于三维空间对称性的著名猜想。

Fourth, Jacob Tsimerman, and the one I find most surprising. His work rests on a tool imported from mathematical logic.第四位,Jacob Tsimerman,也是我觉得最出人意料的一位。他的工作依赖于一个从数理逻辑引入的工具。

Logicians study something called o-minimality, which you can think of as a rule book for tame shapes.逻辑学家研究一种叫做 o-minimality 的东西,你可以把它想成一套关于温顺形状的规则手册。 In a tame setting, the objects you are allowed to define cannot wiggle infinitely often, cannot be pathological, cannot do the horrible things that functions are capable of in full generality.在一个温顺的设定里,你被允许定义的对象不能无限次地摆动,不能是病态的,不能做出函数在完全一般的情形下可能做出的那些糟糕的事情。 Logicians developed this for their own reasons, about what can be defined in a formal language.逻辑学家出于他们自己的原因发展了这套理论,关注的是在一种形式语言中什么是可以被定义的。

Tsimerman, working with Benjamin Bakker and Bruno Klingler, showed that this logical tameness is exactly the right lens for a hard problem in geometry.Tsimerman 与 Benjamin Bakker 和 Bruno Klingler 合作,证明了这种逻辑上的温顺性恰好是审视几何中一个难题的正确视角。 There is an object called a period map, which records how the shape of a geometric family varies as you deform it.有一个叫做周期映射(period map)的对象,它记录了一个几何族的形状在你对其进行形变时如何变化。 Phillip Griffiths conjectured in nineteen seventy that the image of such a map is always an algebraic object, meaning it can be described by polynomial equations, rather than something wilder.Phillip Griffiths 在 1970 年猜想,这样一个映射的像总是一个代数对象,也就是说它可以用多项式方程来描述,而不是某种更狂野的东西。

Their proof works like this. Show the map is definable in a tame logical structure. 他们的证明是这样进行的。先证明这个映射在一个温顺的逻辑结构中是可定义的。 Then show that anything both tame and analytic must be algebraic. The conjecture falls out.然后证明任何既温顺又解析的东西都必定是代数的。猜想便随之得证。

What I like about this is where the tool came from. Nobody built o-minimality to solve problems in Hodge theory.我喜欢这一点的地方在于这个工具的来源。没有人是为了解决 Hodge 理论中的问题而构建 o-minimality 的。 It came from foundations, from people asking what a formal language can express.它来自数学基础,来自那些追问一种形式语言能表达什么的人。 And it turned out to be the right instrument for a completely different room in the building. That, I think, is the thread through all four.结果它却成了这座大厦里一个完全不同房间的合适工具。我想,这正是贯穿这四位的那条线索。

Wang brought geometric and combinatorial thinking into harmonic analysis.Wang 把几何与组合的思维带入了调和分析。 Deng brought techniques from dispersive equations into statistical physics. Tsimerman brought logic into geometry.Deng 把色散方程的技术带入了统计物理。Tsimerman 把逻辑带入了几何。 Pardon showed two separate fields were counting the same thing. One honest caveat about the headlines.Pardon 证明了两个不同的领域在计数同一个东西。关于这些标题,有一点需要诚实地说明。

You will read that these people cracked century old problems out of nowhere. That is not how it went.你会读到说这些人凭空攻克了百年难题。事情并非如此。 Wang and Zahl built directly on a structural insight of Larry Guth's from twenty fourteen, which itself built on work going back decades.Wang 和 Zahl 直接建立在 Larry Guth 于 2014 年提出的一个结构性洞见之上,而那个洞见本身又建立在可追溯几十年的工作之上。 Deng's result extends Lanford's. Every one of these is the last stone in a wall that many people built.Deng 的结果推广了 Lanford 的结果。这其中的每一项,都是许多人共同垒起的一堵墙上的最后一块石头。 The medal goes to a person, because prizes do. The work does not. Next time, something completely different. Thanks for listening.奖章授予个人,因为奖项本就如此。而工作并非如此。下一期,我们将聊一些完全不同的东西。感谢收听。

Check your understanding

Try answering before revealing — these are the points the episode turns on.

1. Besicovitch showed you can rotate a needle using as little area as you like. Why did that force mathematicians to change what they were measuring, rather than simply settling the problem?

Because area turned out to be blind to the thing that matters. The set still has to contain a segment pointing in every one of infinitely many directions, so it is in some sense large, yet its area can be pushed to zero. A measurement that returns zero for every such set cannot distinguish between them or tell you anything about their structure. Dimension in the Hausdorff sense is a finer ruler: it asks how thoroughly a set fills space at every scale, and it can be a fraction. A set can have zero area and still be fully two dimensional. The conjecture then becomes the sharp statement that survives Besicovitch: you can make these sets as thin as you like in area, but never in dimension.

2. Newton's laws are reversible in time, but heat only flows one way. Why is that a genuine paradox rather than just an odd fact, and what does a rigorous derivation of the Boltzmann equation have to do with it?

It is a paradox because the macroscopic behaviour is supposed to be nothing but the microscopic behaviour of many particles. If the underlying rules have no preferred direction of time, then an arrow of time cannot appear from nowhere when you add up many of them, unless something in the passage from small to large introduces it. Boltzmann's equation does have an arrow. So deriving it rigorously from Newtonian collisions is exactly the demand to show where the irreversibility enters and that it is legitimate. The usual answer involves an assumption about particles being statistically uncorrelated before they collide, which stops being exactly true as the system evolves. That is why Lanford could only prove it for a very short time, and why extending it to long times was the hard part.

3. The MNOP conjecture says two different ways of counting curves give the same answer. Why is proving that worth a great deal more than it would be to simply compute both counts and observe they match?

Checking agreement in examples tells you the two definitions coincide on the cases you checked. A proof tells you they must coincide always, which is a statement about the object being counted rather than about the methods. If two frameworks built from completely different foundations, symplectic and algebraic, with different technical machinery, are forced to agree, the natural reading is that both are measuring something intrinsic to the geometry rather than an artifact of how each was set up. It also has a practical consequence: results and techniques become transferable between two communities that previously could not use each other's work.

4. O-minimality was developed by logicians for reasons internal to logic. What made it the right tool for Griffiths' conjecture, and what does the episode suggest is the general lesson?

O-minimality restricts which sets you are allowed to define so that none of them can be pathological, no infinite oscillation, no wild behaviour. The proof strategy exploits a rigidity that follows: an object that is both tame in this logical sense and analytic has no room left to be anything other than algebraic. So showing that period maps are definable in a tame structure converts an open geometric question into an application of that rigidity. The general lesson is about transplantation. Nobody built o-minimality to answer questions in Hodge theory; it came from asking what a formal language can express. Three of the four medals this year reward exactly this move, carrying a tool across a boundary it was not built for.

5. The episode ends by pushing back on the framing that these mathematicians 'cracked century-old problems'. What is the objection, and why does it matter?

The objection is that each result is the last stone in a wall many people built. Wang and Zahl relied on a structural insight of Larry Guth's from 2014, which itself rested on decades of prior work; Deng, Hani and Ma extended Lanford's 1975 theorem rather than starting from nothing. The headline framing compresses a long collaborative effort into a moment of individual genius. It matters because it misdescribes how mathematical progress actually happens, and because prizes, which by their nature go to individuals, quietly reinforce that misdescription. The medal goes to a person; the work does not.

Further reading

  1. 2026 Fields Medals awarded to four of the world's top mathematicians (Simons Foundation)The announcement, with the official citations for all four. Free.
  2. Terence Tao — The three-dimensional Kakeya conjecture, after Wang and ZahlIf you read one thing, read this. Tao explains the structure of the Kakeya proof for mathematicians, but the first few paragraphs are readable by anyone. Free.
  3. Wang & Zahl — Volume estimates for unions of convex sets, and the Kakeya set conjecture in three dimensions (arXiv:2502.17655)The proof itself. 127 pages. Free.
  4. Deng, Hani & Ma — Long time derivation of the Boltzmann equation from hard sphere dynamics (arXiv:2408.07818)The result that removed Lanford's time restriction. The companion paper, arXiv:2503.01800, goes on to derive the fluid equations. Free.
  5. Pardon — Universally counting curves in Calabi–Yau threefolds (arXiv:2308.02948)The MNOP proof. Free.
  6. Bakker, Klingler & Tsimerman — o-minimal GAGA and a conjecture of GriffithsWhere logic meets Hodge theory. The introduction explains the strategy without the machinery. Free.
  7. John Pardon wins the 2026 Fields Medal for work in symplectic geometry (Quanta)Quanta's profile. They are the best in the business at making this stuff readable. Free.
Episode 002

Learning the equation, not the parameter

How a variational autoencoder turned symbolic equation discovery into continuous optimization — and what its answers actually mean

2026-07-29

Distributed hydrological models need millions of parameter values nobody can measure. The field's fix was to learn a function from soil and terrain to parameters instead — but someone still had to guess the form of that function. Feigl, Herrnegger and Schulz train a variational autoencoder on 32 million equations so the search for the right form becomes ordinary continuous optimization. It beats both hand-written transfer functions and a regional LSTM in ungauged basins. The equations it finds contain the tangent of elevation.

🎙 Listen · 12:46
Transcript

Follows the audio as it plays — tap any sentence to jump there.

Today, a paper I think is genuinely important, and a little strange once you look closely at what it produced.今天要讲一篇我认为真正重要的论文,而且当你仔细看它产出的东西时,会觉得有点奇特。 It is by Moritz Feigl, Mathew Herrnegger and Karsten Schulz at BOKU in Vienna, published in Nature Water in February.作者是维也纳 BOKU 大学的 Moritz Feigl、Mathew Herrnegger 和 Karsten Schulz,二月发表在 Nature Water 上。 The title is: Distilling hydrological and land-surface model parameters from physio-geographical properties using text-generating AI.标题是:用生成文本的 AI 从自然地理属性中蒸馏出水文与陆面模型的参数。 I want to spend most of this episode on how it works and why the idea is clever, because the mechanism is the interesting part.这一集我想把大部分时间花在它的工作原理,以及这个想法为什么巧妙上,因为机制才是有意思的部分。

Start with the problem, because if you do not feel the problem the solution looks like a gimmick. You have a distributed hydrological model.先从问题说起,因为如果你感受不到问题,这个解法看起来就像个噱头。你有一个分布式水文模型。

It divides the landscape into a grid and simulates water moving through every cell.它把地表分成一个网格,模拟水流过每一个网格单元。 Each cell needs its own parameter values: how fast water conducts through the soil, how much the canopy intercepts, how deep the roots go.每个单元都需要自己的参数值:水在土壤中传导的速度、冠层拦截了多少、根系扎得多深。 A simple model needs ten to fifty parameters per cell.一个简单的模型每个单元需要十到五十个参数。 Now take the Upper Danube, a hundred thousand square kilometres, at one kilometre resolution.现在拿上多瑙河来说,十万平方公里,以一公里的分辨率。 That is one to five million numbers you must supply before you can run the model even once.那就是在你哪怕只跑一次模型之前,都必须提供的 100 万到 500 万个数字。

Hoshin Gupta, in the commentary that accompanies the paper, puts it in one line.Hoshin Gupta 在随论文发表的评论中用一句话点明了这一点。 Ten parameters and ten thousand grid cells gives you a hundred thousand unknowns, and you are trying to determine them from a single time series of streamflow at the outlet. 十个参数和一万个网格单元,就给了你 10 万个未知量,而你想仅凭出口处一条单一的流量时间序列来确定它们。 The problem is not hard. It is underdetermined.问题不在于难。而在于它是欠定的。 There is not enough information in the data to pin down the answer, and no amount of cleverness in the optimizer fixes that.数据中没有足够的信息来钉住答案,而优化器里再多的巧思也解决不了这一点。

So the field does something smart.于是这个领域做了一件聪明的事。 Instead of asking what is the conductivity in cell number four hundred and twelve, it asks a different question: what function maps soil texture and terrain and vegetation onto conductivity, anywhere.它不去问第 412 号单元里的传导率是多少,而是问一个不同的问题:什么样的函数能把土壤质地、地形和植被,在任何地方都映射到传导率上。 That function is called a transfer function, and the framework built around it in this model is called multiscale parameter regionalization.这个函数叫做传递函数(transfer function),围绕它在这个模型里构建的框架叫做多尺度参数区域化(multiscale parameter regionalization)。 You apply the transfer function at the finest resolution of your soil and terrain data, then aggregate upward to whatever grid you are running on.你在土壤和地形数据的最精细分辨率上应用这个传递函数,然后向上聚合到你实际运行所用的任何网格。

This is a genuinely good move, for two reasons. It collapses millions of unknowns into a handful of coefficients inside a few equations.这是一步真正的好棋,有两个原因。它把数百万个未知量收缩成几个方程里的少数几个系数。 And because the function is applied to the underlying property maps and then upscaled, the model becomes resolution independent.而且因为这个函数是应用在底层的属性图上、再向上升尺度的,模型就变得与分辨率无关。 You can run at half a kilometre or four kilometres and the parameter fields stay consistent.你可以在半公里或四公里的尺度上运行,而参数场保持一致。

But it moves the difficulty rather than removing it. Now somebody has to write down the functional form.但这只是转移了难点,而没有消除它。现在得有人写下那个函数形式。 And as Gupta says, the mathematical form of these relationships is rarely known.而正如 Gupta 所说,这些关系的数学形式很少是已知的。 In practice they are specified by expert judgement and trial and error, with limited theoretical guidance.在实践中,它们是靠专家判断和反复试错来指定的,理论指导十分有限。 Someone decides that conductivity should be an exponential of a weighted sum of sand and clay content, and then we tune the weights.有人决定传导率应该是砂含量和黏土含量加权和的指数函数,然后我们再去调那些权重。 The structure is a guess, frozen in place decades ago, and every subsequent calibration is conditioned on that guess being right.这个结构是一个猜测,几十年前就被固定了下来,而此后每一次率定都以这个猜测是对的为前提。

That is the target. Feigl and colleagues want to learn the structure itself. Here is the difficulty.这就是目标。Feigl 和同事们想要学到结构本身。困难在这里。

The space of possible equations is discrete and combinatorial. Symbols, operators, nesting. You cannot take a gradient in it.可能的方程所构成的空间是离散的、组合式的。符号、算子、嵌套。你没法在里面求梯度。 The classical tool is genetic programming, which mutates and crosses over expression trees, and it is slow and it wanders.经典工具是遗传编程(genetic programming),它对表达式树进行变异和交叉,既慢又漫无方向。

Their move is to make that space continuous. They train a variational autoencoder on equations.他们的做法是把那个空间变成连续的。他们在方程上训练一个 VAE。 Roughly thirty two million of them, sampled from a context free grammar, which is just a formal rule set that generates syntactically valid expressions.大约 3200 万个,从上下文无关文法(context free grammar)中采样得到,这不过是一套生成语法有效表达式的形式化规则集。 The autoencoder learns to compress an equation into a thirty dimensional vector and to decode it back.autoencoder 学习将一个方程压缩成一个 30 维向量,再把它解码还原回来。 After training they throw away the encoder and keep only the decoder.训练完成后,他们丢弃编码器,只保留解码器。 Now any point in that thirty dimensional space decodes into an equation, and searching for a good equation becomes ordinary continuous optimization.如今,那个 30 维空间中的任意一点都能解码成一个方程,寻找好方程也就变成了普通的连续优化问题。

Now the part I think is the real insight, and it is easy to skim past.接下来是我认为真正关键的洞见,而它很容易被一带而过。

A plain autoencoder trained on equation text would organize the latent space by how equations look.一个只在方程文本上训练的普通 autoencoder,会按方程的外观来组织潜空间。 Two expressions that share symbols would land near each other.两个共享符号的表达式会落在彼此附近。 That is almost useless for optimization, because a small change in symbols can be an enormous change in behaviour.这对优化几乎毫无用处,因为符号上的微小改动可能带来行为上的巨大变化。 Swap a plus for a times and the function is unrecognizable. So they make the autoencoder multimodal.把一个加号换成乘号,函数就面目全非了。所以他们把 autoencoder 做成了多模态的。

It is trained to reconstruct two things from the same latent vector: the equation string, and the quantiles of the values that equation actually produces when you feed it the real physiographic data from the study region.它被训练成从同一个潜向量重建两样东西:方程字符串,以及当你把研究区域真实的地文数据(physiographic data)喂给该方程时,它实际产生的那些值的分位数。 Ten percent, twenty percent, and so on, up to ninety.10%、20%,依此类推,一直到 90%。 That second objective forces the latent space to organize by behaviour rather than by appearance.第二个目标迫使潜空间按行为而非按外观来组织。 Nearby points now mean similar output distributions.如今相邻的点意味着相似的输出分布。 And that is what makes gradient free search in this space actually work, because a small step gives you a small change in what the equation does.正是这一点让这个空间里的无梯度搜索真正奏效,因为一小步只会带来方程行为上的一小点变化。

The second property they need is that the space be navigable, meaning you do not spend most of your search in regions that decode to garbage.他们需要的第二个性质是这个空间要可导航,也就是说,你不会把大部分搜索花在那些解码出来是垃圾的区域里。 That is what the variational part buys them.这正是变分(variational)部分为他们换来的东西。 The latent space is regularized toward a standard normal distribution, so the probability mass sits where the valid solutions are.潜空间被正则化,向标准正态分布靠拢,因此概率质量集中在有效解所在的地方。

The rest of the loop is almost mundane, and that is a compliment.循环的其余部分几乎平淡无奇,而这是一句褒奖。 A global search algorithm called shuffled complex evolution proposes a point. The decoder turns it into an equation.一个叫 shuffled complex evolution 的全局搜索算法提出一个点。解码器把它变成一个方程。 That equation becomes a transfer function inside the mesoscale hydrologic model. The model runs over Germany.该方程成为中尺度水文模型(mesoscale hydrologic model)中的一个 transfer function。模型在德国全境运行。 The simulated discharge is scored against observations. Repeat.模拟径流量对照观测值进行评分。如此反复。 They optimize seven transfer functions simultaneously by stacking several latent spaces, covering saturated hydraulic conductivity, saturated water content, field capacity, root fraction and canopy interception.他们通过堆叠多个潜空间,同时优化七个 transfer function,涵盖饱和导水率、饱和含水量、田间持水量、根系比例和冠层截留。

Two structural consequences are worth pausing on. First, the dimension of the optimization problem is fixed.有两个结构性的后果值得停下来说一说。第一,优化问题的维度是固定的。

It is thirty numbers per equation, whether you are modelling Germany or the entire planet.无论你建模的是德国还是整个地球,都是每个方程 30 个数。 Compare that with differentiable parameter learning, where you learn parameter fields directly.把它和可微参数学习(differentiable parameter learning)比较一下,后者直接学习参数场。 That approach requires the model to be differentiable, which excludes most process based models, and the number of unknowns grows with your domain.那种方法要求模型可微,这就排除了大多数基于过程的模型(process based model),而且未知量的数目会随你的求解域增大而增长。 Here, scaling up increases the number of model runs, which parallelize, not the dimension of the search.而在这里,扩大规模增加的是模型运行的次数——这些是可以并行的——而不是搜索的维度。

Second, structural estimation includes variable selection for free.第二,结构估计免费附带了变量选择。 Because the grammar can build equations from any of the available properties, the search decides which ones matter.由于文法可以用任何可用的属性来构建方程,搜索会自行决定哪些属性重要。 Elevation, slope, aspect, bulk density, sand, clay, leaf area index, mean precipitation, mean temperature, temperature range.高程、坡度、坡向、容重、砂粒、黏粒、leaf area index、平均降水、平均气温、气温变幅。 No separate feature importance step.没有单独的特征重要性步骤。 And if you want to impose known physics, you restrict the grammar to a subset of variables and let it search within that constraint.如果你想施加已知的物理约束,就把语法限制在变量的一个子集上,让它在这个约束内搜索。

Now the results, in a prediction in ungauged basins setting, which is the honest test. One hundred sixty two German basins.现在看结果,在无观测流域预测(prediction in ungauged basins)的设定下——这才是诚实的检验。162 个德国流域。 Optimize on fifty of them over six years. Validate on a hundred and twelve different basins over a different, non overlapping period.在其中 50 个上、用 6 年数据优化。在另外 112 个不同流域、一个不重叠的时段上验证。

Median Nash Sutcliffe efficiency on the validation basins.验证流域上的 Nash Sutcliffe 效率中位数。 The default hand written transfer functions, after being calibrated: zero point two six.默认的手写传递函数(transfer function),经过率定后:0.26。 A regional long short term memory ensemble, the current state of the art for ungauged prediction: zero point six two.一个区域性的 LSTM 集成,无观测流域预测当前的最高水平:0.62。 The AI generated transfer functions: zero point seven zero. But the median understates it. Look at the spread.AI 生成的传递函数:0.70。但中位数低估了它。看看它的分布。

The interquartile range for the generated functions is zero point six four to zero point seven seven, which is tight.生成函数的四分位距是 0.64 到 0.77,非常紧凑。 And catastrophic failures, basins where efficiency fell below minus two: seven of them for the default functions, two for the long short term memory, and zero for the generated ones.还有灾难性失败,即效率低于 -2 的流域:默认函数有 7 个,LSTM 有 2 个,生成函数是 0 个。 The reliability improved more than the average did.可靠性的改善比平均值的改善更大。 Running at half a kilometre through four kilometre resolution barely changed anything, which is the resolution independence paying off.在 0.5 公里到 4 公里的分辨率下运行几乎没有任何变化,这正是分辨率无关性(resolution independence)带来的回报。

I want to be careful about the comparison with the long short term memory model, and to their credit so are the authors.我想对与 LSTM 模型的比较保持谨慎,而且值得称赞的是,作者们也是如此。 Fifty basins is a thin training set for that class of model, and the sample excluded very large basins.对那一类模型来说,50 个流域是个偏薄的训练集,而且样本排除了非常大的流域。 This is not deep learning losing to process based modelling in general. It is a statement about this regime.这并不是深度学习总体上输给了基于过程的建模。它只是针对这一特定情形的一个陈述。 The deeper argument the authors make is different and stronger: the long short term memory gives you discharge at a point, while the process model gives you spatially distributed soil moisture and snow, and you can interrogate it.作者提出的更深层论点则不同、也更有力:LSTM 给你的是某个点上的流量,而过程模型给你的是空间分布的土壤湿度和积雪,而且你可以去追问它。

Now the strange part. Here is one of the equations it found, for saturated hydraulic conductivity.现在到了奇怪的部分。这是它找到的其中一个方程,用于饱和水力传导度(saturated hydraulic conductivity)。

Seventy point six nine, times the quantity: log of bulk density, plus leaf area index divided by mean annual temperature, minus fifty four point six seven divided by the quantity tangent of elevation, minus elevation, minus cosine of sand content, minus nineteen point eight seven.70.69 乘以这样一个量:容重的对数,加上 leaf area index 除以年平均气温,减去 54.67 除以(高程的正切减去高程),再减去砂含量的余弦,再减去 19.87。

Tangent of elevation. Cosine of sand percentage. There is no physical reading of that. It is dimensionally incoherent.高程的正切。砂含量百分比的余弦。这没有任何物理上的解读。它在量纲上是不自洽的。 And the paper is admirably direct about it: these are effective, conceptual parameters, they compensate for structural deficiencies and scale mismatches, and they should not be read as true point scale properties.而这篇论文对此坦率得令人钦佩:这些是有效的、概念性的参数,它们补偿了结构缺陷和尺度失配,不应被读作真实的点尺度属性。

And yet.然而。 When they compared the resulting conductivity map against an independent estimate built from a random forest, and the canopy interception map against the GLEAM dataset, the AI generated fields matched the independent data more closely than the hand written transfer functions did.当他们把由此得到的传导度图与一个基于随机森林(random forest)的独立估计做比较,把冠层截留图与 GLEAM 数据集做比较时,AI 生成的场比手写传递函数更贴近独立数据。 So the functions are behaviourally sound while being physically arbitrary in form.所以这些函数在行为上是可靠的,尽管在形式上是物理任意的。

I think the right way to hold this is that interpretable here means inspectable, not derived.我认为看待这一点的正确方式是:这里的可解释指的是可检视,而不是可推导。 You can read the equation, you can plot the field it produces, you can compare it to independent data and argue about it.你可以读这个方程,可以画出它产生的场,可以把它和独立数据比较并就此争论。 That is a real gain over a neural network that emits a parameter field with no expression attached.相比于一个吐出参数场却不附带任何表达式的神经网络,这是实实在在的进步。 It is not the same as having understood something.但它和真正理解了某样东西并不是一回事。

Gupta sees the larger arc, and this is what makes the paper worth your attention beyond hydrology.Gupta 看到了更大的脉络,而这正是让这篇论文超越水文学、值得你关注的原因。 A geoscientific model is a directed graph.一个地球科学模型就是一张有向图。 Nodes are states, links are process equations, and the parameters come from property to parameter relations.节点是状态,连接是过程方程,而参数来自属性到参数的关系。 Every one of those pieces can be written as text.这些组成部分中的每一个都可以写成文本。 So the same generative machinery could in principle propose the process equations, and then the graph structure itself.所以同样的生成机制原则上可以提出过程方程,进而提出图结构本身。 Which is to say, propose scientific hypotheses in symbolic form, and test them directly against data.也就是说,以符号形式提出科学假设,并直接用数据来检验它们。 This paper does the innermost of those three layers. It is the culmination of about eight years of work in that group.这篇论文做的是这三个层次中最内层的那一层。它是那个研究组约八年工作的集大成之作。

The limits are real and the authors state them. One model, one country, temperate. No monsoon, no arid, no glaciers, no permafrost.这些局限是真实存在的,作者们也明确指出了。一个模型,一个国家,温带地区。没有季风,没有干旱区,没有冰川,没有多年冻土。 The equations are conditioned on the data and on the structure of this particular model.这些方程是以数据、以及这个特定模型的结构为条件得出的。 They point at the CAMELS-SPAT dataset as the natural next test. Here is what I would take away.他们指出 CAMELS-SPAT 数据集是顺理成章的下一个检验对象。以下是我想强调的要点。

The contribution is not that AI wrote an equation.其贡献并不在于 AI 写出了一个方程。 It is the reframing: they turned a discrete symbolic search into a continuous one, and they did it by forcing the representation to organize around behaviour rather than syntax.而在于这种重新框定:他们把一个离散的符号搜索转化成了连续的搜索,而他们做到这一点的办法,是迫使表示围绕行为而非语法来组织。 That trick is not specific to hydrology.这个诀窍并不是水文学所特有的。 Anywhere you have a process based model whose functional forms were guessed by experts decades ago, this is now a viable way to ask whether the guesses were any good.任何地方,只要你有一个基于过程的模型,其函数形式是几十年前由专家猜出来的,如今这就是一条可行的途径,去追问那些猜测究竟好不好。

Check your understanding

Try answering before revealing — these are the points the episode turns on.

1. Why is estimating a distributed model's parameters directly an underdetermined problem, and how does a transfer function change the shape of that problem rather than just shrinking it?

Each grid cell needs its own values for ten to fifty parameters, so a modest domain implies hundreds of thousands to millions of unknowns, while the constraint is essentially one discharge time series per basin. No optimizer recovers that. A transfer function replaces the question 'what is the value in each cell' with 'what function maps soil, terrain and vegetation onto this parameter'. The unknowns collapse to a few coefficients inside a shared equation. Crucially it also changes the kind of object being estimated: the function is applied to the underlying high-resolution property maps and then upscaled, so the parameter fields stay consistent when you change model resolution. That resolution independence is not a side effect, it is why the framework survives being scaled.

2. The central trick is making the space of equations continuous. Why is a plain autoencoder trained on equation text not enough, and what specifically fixes it?

A plain autoencoder organizes its latent space by surface form: expressions sharing symbols end up close together. That is nearly useless for search, because swapping one operator can change the function's behaviour completely, so small steps in latent space would produce wild jumps in performance. The fix is a multimodal objective: the same latent vector must reconstruct both the equation string AND the quantiles of the values that equation produces on the real physiographic data. That forces the geometry to organize by behaviour rather than appearance, so neighbouring points behave similarly and gradient-free search becomes meaningful. The variational part adds the second requirement, regularizing the space toward a standard normal so probability mass sits on valid solutions instead of garbage.

3. Why does this method scale to continental or global domains in a way that differentiable parameter learning does not?

The optimization dimensionality is fixed: about thirty numbers per equation, independent of domain size, because what is being searched is the space of functional forms, not the space of parameter values. Enlarging the domain increases only the number of hydrological model runs, which parallelize. Differentiable parameter learning instead estimates the parameter fields themselves, so unknowns grow with the number of grid cells, and it additionally requires the forward model to be differentiable end to end, which excludes most established process-based models. The trade is that you must be able to run the model many times rather than backpropagate through it once.

4. The generated equation for saturated conductivity contains the tangent of elevation and the cosine of sand percentage. Does that invalidate the claim of interpretability?

It falsifies a strong reading of interpretability, not a useful one. Those terms are dimensionally incoherent and have no mechanistic reading. But distributed-model parameters are effective, conceptual quantities that already absorb structural error and scale mismatch, so no one should expect the learned form to be a physical law. What survives is that the function is explicit: you can read it, plot the field it produces, and check it against independent data. They did, and the generated conductivity and interception fields matched an independent random-forest product and GLEAM more closely than the hand-written functions did. So call it inspectable rather than derived — a real gain over a network that emits a parameter field with no expression attached, and not the same as having understood the process.

5. The AI transfer functions beat a regional LSTM (median NSE 0.70 versus 0.62). What are the two strongest reasons not to read this as deep learning losing?

First, the training regime disadvantaged the LSTM and the authors say so: fifty basins is a thin sample for that model class, and very large basins were excluded, so it was being asked to generalize from a set that does not play to its strengths. Second, the comparison is between different products. The LSTM predicts discharge at a gauge; the process model with learned transfer functions produces spatially distributed soil moisture, snow and fluxes that can be interrogated and used in scenarios. The more defensible claim is not that one method is better, but that the learned transfer functions made the process model competitive with the data-driven state of the art while keeping everything the process model is for. The reliability gap is the more striking number anyway: zero catastrophic failures versus two for the LSTM and seven for the default functions.

6. Gupta describes this as the innermost of three layers. What are the other two, and what would have to be true for them to work?

A geoscientific model is a directed graph: nodes are states, links are process equations governing fluxes, and the parameters of those equations come from property-to-parameter relations. This paper generates the third layer, the property-to-parameter equations. The next layer would generate the symbolic process equations themselves, and the outermost would generate the graph structure — which states exist and how they connect. Each is expressible as text, so the same generative machinery applies in principle. What would have to hold is that a latent space can be organized by behaviour for those objects too, which is much harder: the behaviour of a process equation only exists once it is embedded in a working model with the other layers fixed, so the evaluation is far more expensive and far more entangled. That is the same reason this paper optimizes seven transfer functions inside one fixed model rather than searching the model itself.

Further reading

  1. Feigl, Herrnegger & Schulz (2026). Distilling hydrological and land-surface model parameters from physio-geographical properties using text-generating AI. Nature Water 4, 158–168The paper itself. Methods section doubles as a blueprint: model choice, Sobol sensitivity analysis for picking which parameters to learn, the grammar, the VAE, the SCE-UA loop.
  2. Gupta (2026). Deep learning can facilitate physically interpretable geoscientific modelling. Nature Water 4, 118–119The News & Views. Two pages, and the clearest statement of why this matters: the three-layer vision of generating P2P equations, then process equations, then the model graph itself.
  3. Samaniego, Kumar & Attinger (2010). Multiscale parameter regionalization of a grid-based hydrologic model at the mesoscaleThe MPR paper — where the transfer-function-plus-upscaling idea comes from, and why the approach is resolution independent.
  4. Kratzert et al. (2019). Toward improved predictions in ungauged basins: exploiting the power of machine learningThe regional LSTM benchmark they compare against. Worth reading to judge for yourself whether 50 training basins was a fair test.
  5. Klotz, Herrnegger & Schulz (2017). Symbolic regression for the estimation of transfer functions of hydrological modelsWhere this line of work started, eight years earlier — symbolic regression for the same problem, before the latent-space reformulation.
  6. CAMELS-SPAT: a large-sample dataset with spatially distributed forcing and attributesThe dataset the authors name as the natural next test — gridded rather than catchment-aggregated, so it can support distributed models outside Germany.
Episode 001

The death that was never made public

A six-year-old, a bespoke base editor, and the oversight gap it fell through

2026-07-28

In March 2025 a six-year-old girl in Shanghai received a gene-editing therapy built for her single mutation. She died seven days later. The death stayed unpublished for sixteen months, and the paper describing the underlying science left out who paid for it. This episode walks through what happened, what the animal data had already shown, and why the delivery vehicle — not the editor — is the part that killed her.

🎙 Listen · 9:39
Transcript

Follows the audio as it plays — tap any sentence to jump there.

In March of twenty twenty-five, in a hospital in Shanghai, a six-year-old girl received an infusion of trillions of engineered viruses into the fluid around her spinal cord.2025 年 3 月,在上海的一家医院,一个六岁的女孩接受了一次输注,数以万亿计的工程化病毒被注入她脊髓周围的液体中。 The viruses carried a gene editor, built to correct a single wrong letter in her DNA. Seven days later she was dead.这些病毒携带着一个基因编辑器,专门用来纠正她 DNA 中一个错误的字母。七天后,她去世了。 And for sixteen months, almost nobody outside that hospital knew it had happened.而在此后的十六个月里,那家医院之外几乎没有人知道这件事发生过。

The story came out in late July of twenty twenty-six, in a joint investigation by the journal Science and the watchdog Retraction Watch.这件事在 2026 年 7 月下旬曝光,来自《科学》杂志与监督机构 Retraction Watch 的一次联合调查。 I want to walk through it carefully, because there are really three separate failures stacked on top of each other here, and they are not the ones people usually reach for.我想仔细梳理一遍,因为这里其实叠加着三个各自独立的失败,而它们并不是人们通常会想到的那几个。 This is not a story about gene editing being too dangerous to try.这不是一个关于基因编辑太危险、不该尝试的故事。 It is a story about a delivery vehicle with a known toxicity profile, an ethics review that ran ahead of its own safety data, and a family that paid for the privilege.这是一个关于运载工具的故事——它有着已知的毒性特征;关于一次伦理审查——它跑在了自己的安全数据前面;以及关于一个家庭——他们为这份「特权」付了钱。

Start with the child. Science calls her Mei. She had Snijders Blok-Campeau syndrome, a rare neurodevelopmental disorder.从这个孩子说起。《科学》称她为「梅」。她患有 Snijders Blok-Campeau 综合征,一种罕见的神经发育障碍。 In her case it was caused by a mutation in a gene called CHD3, a single base change, one letter of DNA that should have been a C and was a T.在她这个病例中,病因是一个名为 CHD3 的基因发生了突变,一处单碱基改变——DNA 中一个本该是 C 的字母变成了 T。 At six she spoke in simple sentences and ate with training chopsticks.六岁时,她能说简单的句子,用训练筷吃饭。

A single wrong letter is exactly the kind of target that makes base editing look irresistible. Base editing is a refinement of CRISPR.单个错误的字母,正是那种让 base editing 看起来无法抗拒的靶点。base editing 是 CRISPR 的一种改良。 Where classic CRISPR cuts both strands of DNA and lets the cell repair the break, a base editor does not cut.经典的 CRISPR 会切断 DNA 的两条链,再让细胞去修复这个断口,而 base editor 不做切割。 It parks on the target and chemically converts one letter into another. Adenine into guanine, in this case.它停靠在靶点上,通过化学方式把一个字母转换成另一个。在这个例子里,是把腺嘌呤转换成鸟嘌呤。 For a disease caused by one wrong letter, you can draw the fix on a napkin. The hard part was never the editor. The hard part was delivery.对于一种由单个错误字母引起的疾病,你可以在餐巾纸上就把修复方案画出来。难点从来都不是编辑器本身,难点在于递送。

To edit neurons you have to get the machinery into the brain, and the standard vehicle for that is adeno-associated virus, or AAV.要编辑神经元,你必须把这套机器送进大脑,而这方面的标准运载工具是腺相关病毒,即 AAV。 Here the team used AAV9, which can cross into the nervous system. But the base editor was too large to fit inside a single AAV.这里团队使用的是 AAV9,它能够进入神经系统。但 base editor 太大,装不进单个 AAV 里。 So they split it across two separate viral vectors, and injected both into her spinal fluid.于是他们把它拆分到两个独立的病毒载体中,并把两者都注入了她的脊髓液。 Both halves then had to find their way into the same neuron, in the same cell, to reassemble into a working editor.接着,这两个部分必须各自找到进入同一个神经元、同一个细胞的路径,才能重新组装成一个可以工作的编辑器。

Hold on to that detail, because it cuts two ways.记住这个细节,因为它是一把双刃剑。 It is an efficacy problem: the odds of both vectors reaching the same neuron, in enough neurons, to change how a six-year-old brain works, are not good.它是一个疗效问题:两个载体都到达同一个神经元、并且在足够多的神经元里都做到这一点、从而改变一个六岁大脑运作方式的几率,并不乐观。 Several of the experts who later reviewed the case said exactly that. But it is also a dose problem.后来审阅这个病例的几位专家说的正是这一点。但它同时也是一个剂量问题。 When you need two vectors instead of one, and you need them to converge, the natural response is to give more virus.当你需要的是两个载体而不是一个、而且还需要它们汇聚到一起时,自然的应对办法就是给更多的病毒。 And with AAV, dose is the thing that hurts you. Seven days after the infusion, Mei died of thrombotic microangiopathy.而对 AAV 来说,剂量恰恰是伤害你的东西。输注七天后,梅死于血栓性微血管病。

That is a condition where tiny blood clots form throughout the small vessels, chewing up platelets and red blood cells and starving organs, the kidneys especially.这是一种小血管中到处形成微小血栓的病症,它消耗掉血小板和红细胞,使各器官——尤其是肾脏——供血匮乏。 The hospital's own ethics board concluded the death was, in their word, definitely related to the treatment.医院自己的伦理委员会得出结论,用他们的话说,这次死亡与治疗「明确相关」。

Now, here is the part that matters for anyone working in this field. Thrombotic microangiopathy after high-dose AAV is not a freak event.现在,这里有一部分对这个领域的每一位从业者都很重要。高剂量 AAV 后出现血栓性微血管病并非离奇的意外事件。 It is a documented, mechanistically understood complication.它是一种有文献记载、机制上已被理解的并发症。 The virus capsid activates complement, the antibody-driven arm of the immune system, and in a subset of patients that cascade tips over into clotting.病毒衣壳会激活补体,也就是免疫系统中由抗体驱动的那一支,而在一部分患者身上,那条级联反应会失控演变为凝血。 It typically shows up one to two weeks after dosing. It has been reported after approved AAV therapies, not only experimental ones.它通常在给药后一到两周出现。它在已获批的 AAV 疗法之后也有报告,而不只是实验性疗法。 There is a body of literature on it. So the mechanism that killed her was not an unknown unknown.关于它有一批文献。所以,杀死她的那个机制并不是一个「未知的未知」。 It was a known risk of the vehicle, arriving right on schedule. Which brings us to the animal data.这是该载体已知的风险,正如预期般准时出现。这就引出了动物实验数据。

Before the child was dosed, the team ran a toxicology study in monkeys. All four treated animals developed moderate to severe liver damage.在给这个孩子用药之前,团队在猴子身上做了一项毒理学研究。四只接受治疗的动物全部出现了中度至重度的肝损伤。 One animal at the high dose also showed kidney injury.其中一只高剂量组的动物还出现了肾损伤。 Those are exactly the organs you would watch if you were worried about this class of toxicity.如果你担心这类毒性,这些正是你会重点监测的器官。

The hospital ethics committee approved the single-patient trial before it had reviewed the final toxicology report.医院的伦理委员会在审阅最终毒理学报告之前,就批准了这项单一患者试验。

Read that sentence again, because it is the hinge of the whole story. The safety signal existed. It was in the team's own data.把这句话再读一遍,因为它是整个故事的关键所在。安全性信号是存在的。它就在团队自己的数据里。 The body responsible for weighing risk against benefit signed off before it had seen the finished analysis. Then there is the money.负责权衡风险与获益的机构,在看到完整分析之前就已经签字放行。接下来是钱的问题。

According to the investigation, the family paid more than eight hundred thousand US dollars, drawn from their own savings and from relatives, to fund the development of the therapy their daughter received.根据调查,这家人支付了超过 80 万美元,这些钱取自他们自己的积蓄和亲戚,用以资助他们女儿所接受的这种疗法的研发。 Set aside the legal question for a moment and sit with the structural one.先把法律问题放到一边,来仔细想想结构性的问题。 When the family of the patient is also the funder of the research, the pressure to proceed does not sit where it should.当患者的家属同时也是研究的出资方,推动研究继续进行的压力就没有落在它本应在的地方。 Nobody in that room is a disinterested party. In early twenty twenty-six, the team published the preclinical animal work in Nature.那个房间里没有一个人是不带利害关系的旁观者。2026 年初,团队在《自然》上发表了这项临床前动物研究。

The paper described the science.这篇论文描述了其中的科学。 It did not mention the family, or their money, and it did not report that a child had been dosed and had died.它没有提到这家人,也没有提到他们的钱,更没有报告有一个孩子已经接受用药并已死亡。 It referred, in general terms, to challenges in preclinical research and clinical translation.它只是泛泛地提到了临床前研究和临床转化中的挑战。 Nature has since said it was not aware of the issues surrounding the clinical trial when it published.《自然》此后表示,在发表时它并不知晓围绕这项临床试验的种种问题。 The girl's parents have asked the authors to withdraw the paper.这个女孩的父母已经要求作者撤回论文。 Their words, reported by Retraction Watch, were that learning about the missing safeguards has fundamentally changed how they now view the entire project.据 Retraction Watch 报道,他们的原话是,得知那些缺失的安全保障之后,他们如今看待整个项目的方式已经发生了根本性的改变。

So how does a single-patient experiment like this happen without a national regulator ever looking at it?那么,这样一项单一患者的实验,怎么会在国家监管机构从未过目的情况下就得以进行呢? This is the third failure, and it is structural. China runs what is often described as a dual-track system for biomedical research.这是第三重失败,而且是结构性的。中国实行的是一套常被描述为双轨制的生物医学研究体系。 Commercial drug trials go through the national regulator.商业性的药物试验要经过国家监管机构。 But non-commercial, investigator-initiated trials at major hospitals can proceed on the strength of local institutional review, with registration in a national database.但大型医院里非商业性的、由研究者发起的试验,凭借本机构的审查、加上在国家数据库中登记,就可以推进。 No national approval required.无需国家层面的批准。 For a first-in-human gene editing therapy delivered into a child's central nervous system, the entire risk assessment sat inside one institution.对于一项首次用于人体、递送进一个孩子中枢神经系统的基因编辑疗法,整个风险评估都落在一家机构内部。 The same institution that stood to benefit from it working. Afterwards, the hospital paid a fine to the local health authority.而这家机构正是疗法一旦奏效便能从中获益的一方。事后,医院向当地卫生主管部门缴纳了一笔罚款。

The amount was about twenty-four thousand yuan. Around thirty-five hundred US dollars.金额约为 2.4 万元人民币,大约 3500 美元。 The lead investigator, the neuroscientist Zilong Qiu of Shanghai Jiao Tong University, faced no public sanction at the time.主要研究者、上海交通大学的神经科学家仇子龙,当时并未受到任何公开处分。 Neither he, nor the university, nor the hospital responded to the reporters' questions before publication.在报道发表前,他本人、校方以及医院都没有回应记者的提问。 After the investigation appeared, Shanghai Jiao Tong University announced it is now reviewing the case.调查报道出现后,上海交通大学宣布,目前正在对此事进行调查。

I want to be careful about what this story does and does not mean, because the easy reading is the wrong one. 我想谨慎地厘清这个故事意味着什么、又不意味着什么,因为那个轻易得出的解读是错的。

The easy reading is: gene editing killed a child, therefore slow down gene editing. But look at what actually killed her.那个轻易的解读是:基因编辑害死了一个孩子,所以要放慢基因编辑的步伐。但请看看真正害死她的到底是什么。 The editor is not implicated. The delivery vehicle is. And bespoke, single-patient gene editing is not inherently reckless.编辑器本身没有问题,有问题的是递送载体。而为单个患者量身定制的基因编辑,本身并不鲁莽。 One month before Mei was dosed, a baby in Philadelphia named KJ received a base-editing therapy designed for his own private mutation, a metabolic disorder called CPS1 deficiency.在 Mei 接受给药的一个月前,费城一个名叫 KJ 的婴儿接受了一种 base editing 疗法,这种疗法是针对他自己独有的突变设计的,对应的是一种叫做 CPS1 缺乏症的代谢紊乱。 Same underlying technology. Different delivery, lipid nanoparticles to the liver rather than high-dose virus to the brain.底层技术相同。递送方式不同——用脂质纳米颗粒送往肝脏,而不是把高剂量病毒送入大脑。 Different oversight, done under FDA review. Different transparency, published in full including what did not work.监管不同,是在 FDA 审查下进行的。透明度不同,完整发表,包括那些没有奏效的部分。 He went home from the hospital and he is doing well. So the promise is real. That is precisely why the guardrails matter.他出院回了家,现在情况良好。所以这份前景是真实的。而这恰恰是护栏为何重要的原因。

The through-line, then, is not the technology. It is that every safeguard in this case was inside the same building.那么,贯穿始终的主线并不是技术。而是在这个案例中,每一道防线都在同一栋楼里。 The people evaluating the risk, the people delivering the therapy, the people who would author the paper, and the funding, all of it converged on one team, in one institution, with no external body positioned to say wait.评估风险的人、实施疗法的人、将要撰写论文的人,以及资金,全都汇聚到一个团队、一个机构身上,没有任何外部机构处在能够说一句“等等”的位置上。 Under those conditions, the primate liver damage becomes something to work around rather than something to stop for.在这样的条件下,灵长类动物的肝脏损伤就变成了一个需要绕开的问题,而不是一个需要为之叫停的问题。 And when it goes wrong, the same closed loop decides how much of it the world hears about. Which, for sixteen months, was nothing.而当事情出错时,仍是同一个闭环来决定外界能听到多少。而在长达十六个月里,外界听到的是零。

If you take one thing from this episode, make it this. The interesting question is not whether we should edit genes in children.如果你要从这一期节目里记住一件事,那就记住这一点:真正值得追问的问题,不是我们是否应该给儿童编辑基因。 In some cases we clearly should. The question is who, outside the room, has the standing and the information to stop it.在某些情况下我们显然应该这么做。问题在于,在房间之外,谁有资格、又有足够的信息去叫停它。 In this case, no one did.在这个案例中,没有人做到。

The full write-up, with links to the original investigation and to the papers on AAV immune toxicity, is on the episode page.完整的文字稿,连同指向原始调查以及 AAV 免疫毒性相关论文的链接,都在本期节目页面上。 If there is a topic you want covered, there is a feedback link there too. Thanks for listening.如果你有想让我们探讨的话题,那里也有一个反馈链接。感谢收听。

Check your understanding

Try answering before revealing — these are the points the episode turns on.

1. The episode argues the proximate cause of death was the delivery vehicle, not the gene editor. What is the evidence for that distinction?

She died on day seven of thrombotic microangiopathy — a complement-mediated reaction to the AAV capsid that is documented across AAV programmes, including approved ones, and that characteristically appears one to two weeks after dosing. Nothing in the reporting implicates the base-editing chemistry itself. The distinction matters because it changes what the corrective lesson is: this is evidence about vector dose, serotype and route, not about editing precision or off-target effects.

2. Why did splitting the base editor across two AAV9 vectors create an efficacy problem AND a safety problem at the same time?

Efficacy: a functional editor only exists where both halves land in the same neuron, so productive editing scales with the product of two transduction probabilities rather than one — the yield falls off sharply. Safety: the obvious way to compensate is to raise the total capsid dose, and AAV toxicity is dose-driven. So the same design choice simultaneously lowered the expected benefit and raised the expected harm. That is the worst possible direction for a risk-benefit calculation to move.

3. The primate study showed liver damage in all four treated animals. Why is 'the ethics committee approved before reading the final toxicology report' the more serious finding?

Preclinical toxicity signals are ordinary; they exist to be weighed. A signal that is seen can be managed — lower the dose, add monitoring, tighten eligibility, or decline. A signal that arrives after the approval cannot influence the decision at all. The failure is therefore procedural rather than scientific: the one body whose entire function is to weigh risk against benefit acted without its own evidence base.

4. What structural feature made it likely that a bad outcome would go unreported, independent of anyone's intent?

Every function sat inside a single institution: risk assessment, dosing, authorship, and — through the family's payment — funding. No external party held both the standing and the information needed to say stop. The same closed loop that approved the trial also controlled what was disclosed afterwards, and nothing in the publication process forced the preclinical paper to mention that a patient had been dosed and had died.

5. Baby KJ received a bespoke base-editing therapy one month earlier and did well. Name two differences that plausibly changed the risk profile — and one thing that was the same.

Same: a base editor custom-built for one child's private mutation, designed and manufactured on a compressed timeline. Different: delivery by lipid nanoparticle to the liver rather than high-dose dual AAV9 into the central nervous system, which carries a different immunotoxicity profile and is redosable; and oversight by a national regulator under an IND rather than local institutional review alone. A third difference is transparency — that case was published in full, including what did not work.

6. If you were designing the guardrail that would have caught this, where exactly would you put it — and what makes your answer hard?

There is no single clean answer, which is the point. Requiring national review of every investigator-initiated trial would catch it but would also slow the pathway that makes ultra-rare bespoke therapy possible at all. Requiring the ethics committee to sign a completed toxicology package is cheap and would have caught this specific case. Barring patient families from funding the work removes the conflict but also removes the money, since no sponsor develops a therapy for one patient. The defensible minimum is probably: mandatory external review whenever the trial is first-in-human, the funder is the patient, or the delivery vehicle has a known dose-limiting toxicity — and mandatory public registration of outcomes, including deaths.

Further reading

  1. Exclusive: Death of girl in Chinese gene-editing trial was never made publicThe original investigation, by Science with Retraction Watch. Paywalled for some readers.
  2. A couple paid more than $800,000 for a gene-editing therapy for their daughterThe Retraction Watch half of the same investigation — freely readable, and the best account of the funding and the paper.
  3. Brain-directed base editing ends in deathThe clearest technical write-up: the CHD3 mutation, the dual-AAV9 design, the primate toxicology, the cause of death.
  4. Chinese scientist faces probe after girl, 6, dies following experimental gene therapyFollow-up on the institutional response, including the size of the fine.
  5. Thrombotic microangiopathy following systemic AAV administration is dependent on anti-capsid antibodiesThe mechanism behind the cause of death — why high-dose AAV triggers complement-mediated TMA. Open access.
  6. Immune toxicities in AAV gene therapy: an overview for cliniciansBackground on the dose-toxicity relationship that frames this whole case. Open access.
  7. World's first patient treated with personalized CRISPR gene editing therapyThe contrast case: baby KJ, a bespoke base editor delivered by lipid nanoparticle under FDA oversight, published in full.