Databricks just ran a real-world coding test (not some benchmark suite) pitting different agent tools against various models. Chart’s simple: higher on the Y = better success rate, further left on the X = cheaper. Pi with top models is sitting at 80%+ success while burning way fewer tokens than everything else. Been running Pi myself for a bit now — grok-4.5 + DeepSeek-v4 plus my own setup (main agent orchestrating, subagents running concurrent with isolated context). Barely any rework, stupid fast, finally stopped refreshing my phone waiting for results. Main agent + concurrent subagents is clearly the direction. Seeing more tools start shipping this workflow lately.
Databricks just ran a real-world coding test (not some bench… Databricks just ran a real-world coding test (not some bench…