WHEN MUST I START KICKING AND SCREAMING AT YOU THAT IT IS FUCKING HAPPENING
BECAUSE IT IS FUCKING HAPPENING!!!
[day 3/5 - epistemic status: quite an emotional piece, having a bit of a crashout, good post if you are the type of person that needs to be shaken. Did due diligence but the resort to appeals to emotion often.]
EDIT: Post has left intended audience to the point that some have asked what “it is happening” is even referring to. I broadly mean: exponential feedback loops speeding up AI research and capabilities, the technological singularity. This post is not about the labor market, althought it seems plausible that mass automation could be imminent, but that’s really not my wheelhouse.
The METR time horizon benchmark was a very simple idea : take a lot of coding tasks, time humans at completing them and plot average hourly AI success rate at the tasks on the Y axis, and release date on the X-axis1
As of today, it is all but done for and saturated. The company that entirely self-selects for people thinking that AI will be a massive deal has underestimated AI progress.
Having talked about timelines to maybe one too many frontier lab employee now, all of them tell me the same thing.
It’s fucking happening. I can’t tell you anything more than that it’s fucking happening, but it’s fucking happening.
I think you should weigh their evidence very highly. They work for a company whose entire 300 billion dollar evaluation is conditional on keeping a handful of secrets that put them maybe 3 months ahead of the competition. They are way more intimately familiar than us about what makes progress go, what limits are in sight and evidence they have of current capabilities not available to the general public. And unlike CEOs, their job does not depend on hyping up fast timelines.
Quite the contrary, most of the frontier lab employees I talk to believe they won’t have a job in 3 years, because there will barely be jobs in another 3 years.2 It would certainly be way more convenient for them to believe otherwise, but they don’t.
Opus 4.5 was a major update for me and many, I am not writing code anymore. Very little Anthropic employees are writing code anymore. That was almost 3 months ago now. But this realization has not propagated through broader society, at all.3
The amount of people that have used opus 4.6-level capabilities in a coding scaffold, where it really shines, is about 0.04% of the world population.
99.7% of the world’s exposure to AI is getting 5 messages of GPT-5-non-thinking-distill-quantize-2bit-deluxe-instant-fast-spark and the rest is gpt-5 mini.
If you haven’t used them in a while, free models really are utter dogshit, and for the vast majority, it’s all they know. I think the vibe for many is that GPT-5-non-thinking-distill-quantize-2bit-deluxe-instant-fast-spark4 is all AI will ever be. We’re taking off and the masses don’t even know that there’s a rocket5
We evolved to pluck berries of trees and chase animals, and at best discover a couple cool tricks in our lifetime. Progress happening this fast is so insanely unprecedented, do not trust your mind to accurately model the future here, it’s very bad at extrapolating exponentials.
How the FUCK are we supposed to make any anticipatory policy on AI progress when people are THIS badly calibrated on AI progress.
Basically only a subset of the red squares have truly felt whats going on. About 2000 people have felt true, locked-behind-vault frontier capabilities. Almost none of them are lawmakers6, almost none of them people in power7
We seem to be heading to a world where silicon valley will imminently use these technologies to earn billions while a huge swath of SWEs/managers still seriously believe that this is heading nowhere, because they are so jaded by early models.
If you are one of the people that has not seriously considered the transformative impact of AI on the world in the next 5 years. What will convince you. Please draw your rubicon. Do it now. “if AI does this, then I will start taking so and so seriously.”8
Because every step is boring, every step is normal but this is everything but.
God has no bias for normalcy
There are reasons to believe we are at, not only in an exponential scenario, but a superexponential.
The amount of lab compute is growing 4x every year which is a main bottleneck for progress
Epoch projects that the by the end of the decade, we will have a model that exceeds GPT-4 in scale to the same degree that GPT-4 eclipses GPT-2 in training compute.9
Algorithmic progress/efficiency has been growing at 4x every year
Coding models are now good enough to radically speed up the development of AI progress10, the average Anthropic employee reports opus 4.6 lets them do in one day what used to take 2 and a half. some say it lets them do in a day what used to take a week. This was not the case previously, the models we see now have not benefitted a lot from this speedup, research lags behind on production.
These things stack multiplicatively. (yes, like in balatro)
Forethought projected all of this out into the future, with 2 timelines, one where there are no feedback loops.
Why leave out the feedback loops? Because the immediate future looked way too weird with them.
I am telling you that the opus 4.6 system card is literally telling you that the people at these companies are directly observing the feedback loop
Oh my sigmoid
I have looked at exclusively coding for now, but coding is a window into everything else. Making the model better at things is easier if you can just give it a task and make it do the thing by itself.
Technological growth generally follows a sigmoid curve, which is the main argument against continued AI progress. This generally happens because low hanging fruit gets picked, and getting ideas for improvement becomes harder and harder
But generally, technological growth has not resulted in 7x self reported speedups on improving said technological growth, in a feedback loop where the path is now exceedingly clear.
I agree that “this time is different“ claims should always come with great epistemic humility. But AI progress is climbing on like 8 different curves simultaneously, all inputs are only increasing, thinking they will all converge on a sigmoid any time soon is quite silly.
There’s no strong reason to believe that making a coding model good at 1 year time horizons is significantly harder than making it good at 16 hour time horizons. 16 hours is a sufficiently long time that it requires all the things like planning, execution, realizing when what you did was stupid, etc… Extending this to years seems not qualitatively different. And that any barrier preventing this would have shown up already.
In even in the most bearish of bearish of cases there is still a shit ton of progress to be had. You can’t seriously make the claim that we have exhausted literally every RL environment. You can’t seriously make the claim that more compute will not impact progress. You can’t seriously make the claim that current and imminent models will not speed up research. You can’t seriously make the claim that the growth of AI researchers will abruptly grind to a halt
Remember, things will never get worse, it will never be 2022 again. We have the recipe for current models, and we can just see if a new idea makes them better and refuse to implement it if it wont. Progress will never regress, barring some apocalyptic scenario. Opus 4.6 is the worst model you will ever talk to for the rest of your life.
So I ask you: how many more private benchmarks must fall before you open your eyes?
The METR plot is saturated, datacenter buildout is only accelerating, Dario won’t pause, Demis isn’t contacting us. WE MIGHT BE ALONE, IT MIGHT JUST BE YOU AND ME. AND WELL, 50 SCRAWNY EA WHITE GUYS IN A BOARD MEETING ROOM DECIDING THE FUTURE OF HUMANITY! BUT THAT'S OKAY! BECAUSE DO YOU REALLY NEED ANYONE ELSE?!?!?11
If this post has left you feeling hopeless, or that you want to help in some sort of way, please check out this, this, this & maybe even this
I’m cognisant of the dangers of believing in things that the wider population thinks is a little outrageous, The bay and its people are obviously having some sort of influence on me, might make a more epistemically grounded piece about this later
Nuance: hill climbable, majority of the tasks is private. The confidence intervals are pretty wide, less so for 80%, where the trend is less agressive and follows an exponential rather than a superexponential. Anthropic explicitly wants to get good at coding, so hillclimbing this is probably doubly effective for them. A perceived capabilties gap between benchmarks and real world usage holds, although it is narrowing. That being said, the point of the article still holds
aka, whatever is released to the public, but 3-6 months ahead
sources on the infographic: Global AI adoption rate (16.3%) from Microsoft AI Economy Institute H2 2025 survey. Paid subscribers estimated from OpenAI COO Brad Lightcap's statement of "more than 10 million" ChatGPT paying users, plus an estimated ~50% additional across Claude Pro, Gemini Advanced, Copilot Pro, etc. (~15-25M total). Coding scaffold users estimated from Cursor's 1M+ reported users (DevGraphiq), plus Claude Code, Codex CLI, Windsurf, and similar tools (~2-5M total). World population: 8.1B (UN, 2024). All figures are order-of-magnitude estimates; the precise boundaries between tiers are fuzzy but the ratios are robust. Vibe research by claude
comedic effect
worse of all, this gap will only widen as exponential progress does its exponential thing, and pricing becomes more and more expensive so serving free models becomes harder and harder.
shoutout pete buttigieg visiting the bay and getting pilled on claude code
well, maybe financially, soon
and I would invite you to ponder if it has already happened
if you don’t know what this means, GPT-2 can barely coherently form sentences
“Productivity uplift estimates ranged from 30% to 700%, with a mean of 152% and median of 100%“ page 184









You forgot to make a case for the point you are trying to make. You just got more agitated and sloppy about evidence as the article progressed.
I like the articles you write where you can tell the cortisol is building up the further down you read