I think that may be quite an unusual situation that in the past, maybe in the dot-com bubble, maybe people were talking about the railroad rush and stuff like this, we did not see this bifurcation. So I think, yeah, I've been thinking about this more, and I think the situation is getting more and more interesting.
Fascinating. So you alluded to some of your predictions or extrapolations for '26 and '27. Do you want to unpack that? You had three of those.
Maybe calling it my prediction is giving myself too much credit. I will just say this: if you look, for example, at METR eval and you very naively extrapolate the linear fit, that's what you would expect to happen. And so I'm just going to be humble and say, most of the time, I'm not going to be smarter than statistical models, statistical extrapolation of past trends that have been very consistent. So I'm just going to be very humble. Despite what I might know about research and what's happening, probably the most likely, the best prediction I can make is actually just follow that data, that extrapolation, and see where it's going to take us.
And yeah, in that case, if you roll this out, if you look at other benchmarks, I think we would have something like next year, maybe the models will be able to work on their own for a whole day's worth of tasks. If you think of software, you might say, like, implement this entire feature, build out this entire set of the app. If you think of knowledge work, maybe do a whole research report, this kind of scale. The reason I think task length specifically is interesting is because that's what allows you to delegate more and more work to language models, to agents.
Even if you have a very clever model, but if it needs feedback or interaction with you very often, then it really limits what you can delegate to it. If you need to talk to it every 10 minutes, right? Versus if you have something that can go for hours at a time, obviously, right? Then you cannot just have one copy of it. You can have a whole team that you delegate tasks to and manage. And so I think that's why it's really critical that the models are actually smart enough, the agents are smart enough, to work on their own, to correct their own errors, to iterate, because that's really what allows you to delegate.
Indeed. Task length and time to complete as the metric for progress. So by mid-2026, you mentioned agents can work all day autonomously. Late 2026, at least one model matches industry experts across many occupations. And then by 2027, models frequently outperform experts on many tasks. So it's more time running and then generalization across the economy. And you mentioned GDPval, the OpenAI metric, as a benchmark to already see the progress towards multiple professions.
Yeah, I think GDPval is a super cool evaluation from OpenAI, where they collected a lot of real-world tasks from real domain experts to make sure it is actually representative of what you might do in the economy. And then they evaluated a lot of models on those tasks. They compared them against real experts' performance to give us a really good indication of how close, how far are we from having significant economic impact. So I think that's a super cool evaluation.
The sort of obvious question is that GDPval and METR are carefully designed benchmarks. How do they predict production value once you add compliance, liability, messy data, the messy world, tool friction, and all the things?
So I think messiness and task length, time duration that you're able to work independently, are very similar or very correlated. So I think that's why it's interesting that METR tries to measure how long can the model go on its own. Because if you think about how do you come up with a task that takes a human eight hours, 16 hours, you will have to include all this messiness and all this real-worldness to even be able to measure it. But I think ultimately, to go further, we really want benchmarks.
We really want evaluations that come from the actual users, whether it's the industry, whether it's private users, because that's what ultimately matters. Is the model helpful to you? Do you get something out of it? Does it do your paperwork, help you write something, fix your code, help you study? I think that's the real proof. If you release a new model, do people start using it more?