If you would look at multilingual, or 60-something percent for just Python, getting that to actually a number, I think 90% is going to be the minimum bar for broad adoption and saying, this is now established technology and we need to look for the next big thing. That is going to be the challenge of the years to come.
Yeah, let's double-click on that for a few minutes. The launch of that agent mode was a major announcement, a major new step in the history of GitHub, that just happened at Build 2025 a few weeks ago. What does it do?
The coding agent, the way it works is that you can just give it a task: coding tasks, documentation, generating test cases, or simple things like, find all the bugs in my codebase. The agent then goes off in the cloud and spins up a virtual machine, checks out the repository, installs other tools, and solves that task for you. And the magic here is that, in the meantime, you can keep working on your part of the codebase, on a different task, a different issue, on your local machine.
And so effectively, the coding agent is like a new member of your team that can take on certain tasks. And you can obviously not only assign one task to one coding agent, you can assign 10 tasks to 10 versions of that coding agent, and they can all run in parallel. And when they're done with their work, they submit a pull request exactly like one of your human team members would do. And then they alert you and say, hey, this pull request is ready for review.
And then you go in and you review the code, and you can comment on it, and Copilot will pick up those comments and keep iterating. And so, if you don't like the code, I tested it yesterday with one of my hobby projects, and I realized the README is completely outdated because I didn't spend time on writing a README for myself, and just told the coding agent, look at the codebase and write a new README. And then it wrote a README, and I reviewed the README and looked at the headline. It changed the title of the README from README to something else.
And so I gave it feedback and said, hey, I think the README should always be called README, and so revert that change. And it just did that, right? And so that's the most simple explanation that I can give of the coding agent. There is nuance here, because there are going to be scenarios where now you want to take that code generated by the coding agent and check it out on your local machine. And you can do that with the GitHub command line and then keep working in VS Code.
And of course, Copilot agent mode is also available for you there. And so you can then use the agent in synchronous mode on your local machine. And so that's, if you stay with the analogy of your team, that's like you pulled that team member that submitted the pull request onto your desk or into a Zoom call. And now you're working together on that codebase. But now your focus is on that task in synchronous mode while the coding agent works asynchronously.
And so this continuous spectrum of: I can assign tasks, test generation, bug fixes, security vulnerabilities, bootstrapping a new feature to the coding agent; I can take that codebase into my local IDE; or I can keep working just as I'm used to, all these things in parallel, that, I think, is going to be the future of software development. And then the skill of the developer, probably, to know how to describe the task in such a way that the agent can do the job with almost no additional revisions needed.
Because every time you have to review code and correct the agent and wait another 10 minutes, you're interrupting your flow again, and you have to do something else in those 10 minutes. And jumping between different— we're bad at multitasking, and we're bad at context switching. And most developers do not want to work on 30 different things throughout the day. That's stressful and ultimately doesn't lead to great products.
And the agent works with prompts, right? It's vibe coding, where you describe what it is that you want it to do and it does it. Is that right?
It's prompts when you use it within the IDE. Although even there, you can start with the brainstorming cycle first. And so that, for example, works really great with Claude Sonnet or Claude Opus. And so you can first ask it, how would I build this? And what's the system design for this feature, for example? And then have it first write a Markdown file with bullets, doing the engineering together with the model. And then you take the first task of that and feed that into the agent mode to write the code.
For the coding agent, because it sits on the GitHub platform, the starting point is an issue. And the issue can be the description that comes from your product manager or from a user of your open-source project. But it's also all the comments in the issue, attached images, file references, or web pages. And of course, the coding agent and agent mode both can use MCP servers and tools. And so you can then further connect into additional contexts.
But yeah, fundamentally, it's a prompt. It's just that the prompt, when it's no longer just one input field with three lines of code, it can be a long description, specification from a product manager, just as they would write for a human developer. And again, the product manager then is the one that needs to learn how much do I need to decompose the big problem into a small building block.
And when you mentioned 30% or 40% success, is that across the board? Or from your experience using the product and early feedback from users, is the agent particularly good at certain tasks versus other tasks? If I'm a GitHub user and I want to play with the product, what should I do first?
The 30%, I was referencing the SWE-bench, S-W-E benchmark that was developed by a number of researchers and originally started as Python only. And it's 2,000 issue-pull request pairs out of a dozen Python open-source repositories, but was recently expanded to other languages. And so in Python, I think the best benchmarks are 60% to 70%, depending on what model-agent combination you take. But if you look at multilingual, we are in the 20% to 30% range for these benchmarks. So that gives you an idea of real-life issues in open-source projects and the corresponding pull requests.
How good is the agent at solving that existing issue compared to the original solution? For our coding agent, the way to approach this for developers is to just go and try it out. Because the worst-case scenario here is that you get a pull request that is so far off from what you would build yourself that you close the pull request without merging it. But then that gave you a learning cycle between the description that you gave it and the code it generated.
And then there's multiple ways you can approach that. If you don't want to throw it away, one is to provide custom instructions in a file within the repository to give the agent more context of what you expect it to do. So it's kind of like a how-to that you provide to the agent in the same way that if you would hire a developer into your team, or you would instruct people in open source to contribute back to your project, you would also give them: this is how we expect the coding standards, the libraries, how we're writing unit tests.
So giving the agent more instructions often leads to better outcomes. Scoping it down to a smaller change often leads to better outcomes and also to smaller pull requests, which is the best practice. And over time, the models will get better. And then three days later, Anthropic came with Claude Sonnet 4. And so now the coding agent is running on the 4 model. And of course, these models will get better, the context for the model will get better, the tool calls will be better.
And so I think the key thing to always remember is whatever your experience was with AI a month ago is probably no longer the state of the art.