This [very hard benchmark](https://github.com/browser-use/benchmark) targets the hardest browser tasks. On easier tasks, even smaller models can achieve very high success rates. Results shown are from a 60-task subset of BU Bench V2.
## Integrations, hosting, custom tools, MCP, and more on our [Docs ↗](https://docs.browser-use.com)