Okay, so this is pretty spectacular as not just an announcement, but as progress. So to unpack it, you're basically saying that the Vercel AI Cloud is making DevOps redundant. I mean, is that a self-healing system?
So there's two parts I could call infrastructure and code. Infrastructure, we've been automating for the past 10 years. And of course, it's all automated. It's on autopilot, it's autonomous, et cetera. What's been harder for people is the operations of that, right? For example, someone attacks your website. I'm talking about a particular targeted attack. We host a really, really popular sneaker website on Vercel. It constantly has bots trying to do credit card stuffing attacks, like trying out stolen credit cards.
They're trying to scrape their inventory, and they're anomalous in nature. Let's say that, 9 to 5 business hours, you're operating your online store. You kind of know for the U.S. what your hot hours are, what your peak traffic looks like. It's really easy, and Vercel already does this, to detect an anomaly. But what we're doing now is we're deploying the Vercel Agent to investigate that anomaly. We're going to tell you what the recommended action is.
Maybe the attack originates from a country that you don't even do business in. It's like, why did traffic from Argentina—just to throw myself under the bus, I'm from Argentina, because I always say, what country? I don't want to say a country that's going to get people upset.
—increase by 100x? I don't want to give this problem to the website, to the developer, to the website builder, et cetera. And so Vercel Agent is going to tell you, you've never received traffic from Argentina. This is 100x. We are going to actually—this is really cool. Obviously, we showed this demo on stage. Because we understand the network so well, we can decode who the user agents are. Are they trying to fake their user agent? Are they trying to do spoofing?
And so we're going to give you all of the facts, first of all, and then we're going to give you the decision. And so you're still in the loop. You make the decision. But imagine the other world, a world of a more traditional cloud where maybe there's some anomaly stuff, but typically an operator is going to get paged with zero context. Literally, the pager goes out. I don't know if you heard this story from Patrick Collison, the CEO of Stripe.
In the early days of Stripe, he was holding the pager, and it was a duck sound because he was trying to make it as stress-free as possible. And he had this really funny tweet of, like, anytime he's walking around in a lake, in a new city, whatever, he hears ducks, his cortisol levels go up automatically because it's the pager sound. But this is the real world of the cloud today. You get paged, you get your ducks, and then you're like, okay, what do I do?
All right, begin the investigation process. Maybe an hour later, if you have someone that's really on the ball, they kind of have an idea of what's going on. So we believe that the AI cloud will be self-healing, as well as completely autonomous from an infrastructure point of view. And we believe that agents and humans will collaborate. So it's not going to be just, like, off-the-rails autonomous. It's going to be with you in the loop. And that's why the pull request is such an important tool here, where we can create code changes.
By and large, the cloud has, I would say, like three—the dark side of the cloud would be three things. Number one is what I talked about. The internet is a very nasty place. Attacks constantly happen. And that's why we've been investing so much in our firewall capabilities, the DDoS mitigation capabilities, et cetera. The other one is stuff goes down. So, you saw the Google Cloud outage the other day for hours. A lot of the cloud was down.
So I'll give you an example. Vercel didn't go down, and this is not me throwing shade at Google, but if you were using a Google Cloud service in Vercel, your endpoint is now down because you took a dependency on Google Cloud. Luckily for us, a lot of our customers did not. We actually looked at the numbers, but there were quite a few customers that had increased 500 errors. This is an error that happens when your website is crashing, but many times your website crashes because of somebody else's stuff crashing.
Vercel Agent will also help you investigate and repair that. Attributing the cause is so important and liberating and empowering. Because if I get an investigation from Vercel Agent that says the reason that you have an elevated error rate is because this outgoing host is down, knowledge is power. I can now—I know what phone to pick up and who to yell at, or I could help you fail over that system. So that's the second category. It's like, when things break, we'll help you.
The other one is optimization. Vercel helps you make the web a lot faster, but you still have—I always tell people, these are Turing-complete systems. We don't yet have the quantum computers that would allow us to simulate every possible future. Things can still get slow in ways that it was hard for the developer or agent to anticipate. And so we can help you. We can emit pull requests that optimize your systems. We can say, look, big e-commerce website, this image that you're rendering from a mobile phone is just awfully slow.
And this is the stuff that I was telling you, like, I would do door to door. I literally open a website and can immediately spot, yeah, that image you're not prioritizing and optimizing correctly. So the Vercel Agent will not just tell you, oh, you're slow, loser, you have a 50 score. It's actually going to be a part of the solution. So that's kind of like the—on the autonomy side. The other thing that an AI cloud needs is services and SDKs purpose-built for AI.
So we announced the AI Gateway in beta. We announced the Vercel Sandbox. Obviously, we announced the Vercel Agent. And we announced improvements to Fluid, which is our compute platform, that make it better and faster and way more cost-efficient to host things like MCP servers and AI workloads.
And Fluid itself is a reasonably recent launch as well, right? That was a few months ago.
Yeah, it's very recent. What the AI transition has done is that the profile of workloads that run in the cloud has changed dramatically. I cannot emphasize enough just how different these programs are. By and large, the big difference is you can call it a transition from frontend to backend. In the past, everything was pixels. They need to render really fast, and you could have short bursts of compute. Amazon invented the 100-millisecond rule. It's a very important rule for the world to know.
For each 100 milliseconds that Amazon takes longer to load, they would lose 1% in conversions. And so we optimize all of our systems around that world of pixels. The world of tokens, I don't know if you've used the new model by OpenAI, o3 Pro.
It can think for 15 minutes. What compute platform was optimized for a request comes in, something gets launched, and it cooks for 15 minutes? Basically none except for Fluid. So what we announced is Fluid now only charges—this is actually only a billing model change. We already did the sort of, like, compute optimizations for running longer, et cetera. But we changed the billing model. If you're waiting for 15 minutes for OpenAI to respond and you do a tiny bit of compute, maybe to transform it into HTML or transform it into an MCP response, or call other systems.
We're only gonna charge you for the actual CPU cycles that you use. Vercel started out with pay for what you use, but you're actually paying for the idle time.
It's kind of a white lie. You're paying for what you allocated, and you most likely overallocated the heck out of your system.
With Fluid Compute, you're actually paying exactly for what you use, in the sense of what you compute. So it's a really, really, really high-efficiency CPU that complements your GPU. So you're still gonna call models like Anthropic's and OpenAI's and Mistral's, et cetera. But you need to do the computation on top of them. And so that's where Fluid really shines.
Fascinating. So it's a reinvention of serverless for the world of AI infrastructure.
In particular around response time, idle time. Okay. What about Sandbox?
Yeah. So one of the fascinating things in the world of agents is that Fluid needs to run the code that humans wrote to power your applications. But what's happening is that increasingly more and more code is being created by agents on demand. A good example is when you're working on things like deep research. Let's say you're a big pharmaceutical company or a bank and you want to create your own deep research system. So first of all, you're going to need this sort of high-efficiency compute like Fluid.
You're going to need, obviously, some of the foundational infrastructure we talked about, like CDNs, firewalls, et cetera. But now you need a new thing that actually doesn't quite exist in the world, which is you need a sandbox, meaning a secure place for compute to run that is generated by the model. So when agents are doing research, they might decide, hey, I need to run a quick Python script to compute some numbers or to produce a visualization of the data, or to make up my mind or whatever.
So you can think of Sandbox as the Amazon EC2 of AI. It's not for the code that your developers wrote; it's for the code that gets emitted by the LLMs. And it allows you to create these incredible new products. I mean, you could create your own v0 or Lovable with this Sandbox primitive, but also you can create systems where the code runs behind the scenes in the service of an outcome, like a deep research report.
In a lot of what you've described, both for the AI Cloud and v0, it seems that you have a constant data flywheel going where you constantly learn and optimize. Is that how you think about it? And to which extent have you formalized it into a product or a part of the product?
Yeah, I think part of the AI Cloud is going to be what we build. Like, we build Fluid, we build the Sandbox, we build the AI Gateway, but also the rich ecosystem of integrations. So for evals, we use a product called Braintrust, for example. And we want to make sure that every product that we don't build but you still need in the AI world is in our marketplace. Another really cool product that I think most people building agents are going to need is you need to give your agent the ability to run code securely, like Sandbox.
You also need to give it a web browser, which is also another kind of sandbox in a way. If you're an agent, you need to be able to browse the web. And so you can expect integrations like Browserbase and Browser Use. They're building browser infrastructure for agents. So the things that we don't build, you're going to have one click away as part of the Vercel Marketplace. Evals—I just cannot emphasize enough how important that data flywheel is.
I think in the old world, there are companies that are more data-driven, there are companies that are less data-driven, right? Like Apple famously has always been a vibes company where their designers sit down, they look at the product, how it feels. It's a very demo culture. We do a lot of that at Vercel. But famously, other companies like Facebook, for example, are very data-driven. Oh, the user stopped at an ad for less than 10 milliseconds. We should create another ad. We need to optimize this.
And if you add seven friends, you're going to be a long-term engaged user and whatever, that kind of thinking. In the AI world, it's not optional to have that data-driven approach where every data point you need to synthesize, you need to pay attention to how users are responding to generations, what feedback they're giving, what's the error rate. Satya was just talking about the acceptance rate of when you propose something. If you're building an agent that proposes changes, you want to monitor very closely the acceptance rate of the things that the agent proposes.