I did a talk at Data Driven about ontologies, whatever, a year and a half ago or something. And I think there are a lot of people who think about ontologies in their own data, in their own repo. And those people who are actively thinking not just, how do I prompt a model, but how do I build a system? They start to build really good intuitions.
So make sure to cover it before the end of the conversation. What is it exactly that you're open-sourcing with Braintrust? Walk us through the project, where people find it, the genesis of it. Why are you partnering with Braintrust specifically on this?
If we go back to the idea of behaviors, right, the idea is that you can actually write down in Markdown what is a— it is actually both a spec and a rubric. We call these specs, and there were some people who asked, isn't this a rubric? And it is. It's both. The reason it's both is because it is not just used to grade or potentially reward the agent. It's also used to align the humans. I think that is an underrated point, in that how you want the agent to behave, as we talked about earlier, is actually a subjective question.
And so internally, you need to build processes to all agree on, hey, this is the product, right? How do you want the agent to behave? Going back to the example about the fast PowerPoint verification. And so you want a standard to kind of write that down. And the project kind of came about because I was actually having coffee with the CEO of Braintrust, Ankur. And I forget why, honestly, but I started telling him, I was explaining this concept to him.
I was talking about this because we were doing this internally, and I thought it was very cool. And he got pretty excited about it. And one thing that I had internally that I'm trying to think about is we, I think, do a lot of really cutting-edge work, but it's not something we talk about much because, to be honest, we're working all the time.
Yeah. Right before we started recording, you were showing some internal Slacks between your co-founder Matt and yourself. And if I may disclose them, Matt was sending you a Slack at 4:00 a.m. and—
Yeah. And that was last night. So it was a Sunday night as we were recording this, and you showed how you were replying to that Slack at 6:00 a.m. So yes, 9-9-6 in full action amongst the co-founders of Basis.
There's no 9-9-6. For Matt and me, it's 24/7. For the rest of the company, people work hard, but it's definitely not a 9-9-6. And so we wanted to talk about it more and just share what we're doing. And we don't have a lot of resources to blast out to people. And so we were talking. I was like, well, I actually think this could be really useful for Braintrust and honestly the whole industry, because if you have a standard that could be something that people define, it can get automatically slurped up into observability platforms, monitoring platforms.
And for people maybe who are less advanced, it could also have out-of-the-box judges or ways to define, hey, here are the behaviors, and you don't have to configure your own judge. You can actually get it to judge it for you and see the results. And so he got pretty excited about that. And so that's where the collaboration came from. So that's what the open-source repo has. It has a couple small examples. It has an example judge that you can use.
It has examples of actually written behaviors that you can leverage and sort of build your own. And I think it is useful to think about how to adopt this standard, but I think it is also, maybe more importantly, thinking about how to adopt the mindset of not thinking that an agent operating over 10 hours is a black box. It's not. It has a lot of data, and you're probably doing a disservice to your customers if you don't understand how it's going about the work.
And what would you want people to do with this open-source project? Presumably contribute to it, use it for their own purposes. How does this become an industry standard?
Yeah, it's a great question. I don't actually know. I think the coolest thing would be for people to contribute ideas to it. I think there's a lot of work left to do. I think it's just the beginning. I mentioned a couple things earlier, but there's so much to do around, one, how you make good judges. Two, how do you properly label and dissect trajectories to make it easier for judges to understand? Because at the current level of expense, you couldn't run this in all of production, for example, because you're running judges on every single trajectory.
But there's a lot that can be done, I think, with building out the work that sits on top of the behaviors. And I think also just seeing, we purposely tried to make the standard relatively flexible, similar to skills, where it is just Markdown. There's not an overfit, hey, you need to have these exact five words. You can make it very broad and you can make it very specific, as long as it is still self-contained to the point that a judge could look at the behavior and actually know, was the condition for it to be exhibited met?
And if so, did it get exhibited or did it not get exhibited? And as long as it has that, there's a lot of leeway there. And so I think we wanted that to be flexible.
So what else should AI builders think about as they build autonomous, long-horizon agents? So we talked about judges, we talked about behavior. You just mentioned ontology, which in your talk at Data Driven NYC, you had mentioned as well as a world for agents to live in. Where does that fit in the picture?
Yeah, they're super important. A lot of people, when they think about agents, their mental model always goes to coding agents because that's everyone's experience with, at least the people who probably listen to this podcast. People think about coding agents a lot. Coding agents are interesting because they obviously have a harness that they get shipped in. They have certain tools, they have certain behaviors encoded in them in the context, right? Codex, by the way, is open source.
I highly recommend people go look at the open-source repo, but they don't control their runtime training data, right? Because their runtime training data, which goes back to my analogy earlier, which is your context, is actually the repo they're working on. And so you could have Codex, the same agent, quote-unquote, on one codebase perform somewhat well, and then another codebase perform spectacularly because that codebase has much better runtime training data, right? It has potentially good skills or good context about how to operate in the codebase because there aren't contradictions and confusions, whatever it is.
And so the ontology of your codebase, it always mattered for engineers. It matters just as much, if not more, for really good agents over time. That actually, I think, goes up another level if you're thinking about non-coding agents, because in coding, you don't own the runtime training data. In non-coding, you do, right? Most of the data, quote-unquote, that an agent sees when it's a Basis agent, Basis owns, right? It's training data that we have to ensure works really well.
And again, when I say training data, I mean, like, effectively handwritten context, right? Or things that are part of your broader progressive disclosure. Some could be handwritten, some could not be, whatever it is. And because the agent is always starting from scratch, designing that ontology in a way that is ergonomic for the agent is super key to building something that's long-running. That's both, like, the data that is maybe static, like skills and whatnot that all the agents have. But also, once you get to really long horizons, if you're talking about a stateful agent that's maybe operating over days to months, now suddenly you have an ontology of, going back to the Memento example, information the agent's left for itself, which, if you're operating for maybe a couple hours, could be a couple notes.
If you're operating for months, you're talking folders, right? And you have so much knowledge and context that describes the lived experiences of the agent that suddenly this new agent—well, new—that is standing up with effectively very compacted context has to get it back into the state of mind of its entire lived experience. And so the ontology designed to make it easy for it to do that, and the behaviors you encode that properly ensure it's doing that well, are sort of the key to making it work really well.