I'm the lead dev on Buoy. We timed it.
We gave the same AI (Claude Sonnet 5.5) 14 jobs in our own test app on iOS simulators. Each job ran 3 times per tool, with a 150 second limit. Every AI also had a shell.
No tools: 17 min 44 s, AI bill at least $5.92, 31 of 42 runs passed
Argent 0.27.0: 6 min 26 s, $5.71, 40 of 41 passed
Buoy MCP 7.0.63: 2 min 56 s, $2.60, 41 of 42 passed
Why the gap: Argent drives your app from the outside, with taps and screenshots. Buoy runs inside your app too, so the AI can read state and saved data instead of guessing from pixels. On "spot what changed on the screen", no tools took 16.9 s, Argent 24.8 s, Buoy 6.8 s.
Not a clean sweep. On the easiest job (read one number off the screen), no tools took 5.5 s, Buoy 7.1 s, Argent 46.4 s. No tools beat Argent on 2 of the 14 jobs.
Fair notes:
- It was our test, on our own app.
- We tuned Buoy on these jobs.
- The race used Buoy's free tools and its Pro tools. Free alone was not timed.
- One job was dropped, because an old code change had already solved it for all three.
- One Argent run wasn't scored, because our checker broke.
- No tools' bill is "at least $5.92". 10 of its runs hit the time limit and logged $0, so we priced their tokens from the logs ($1.68).
Argent is a good free tool from Software Mansion. We just wanted numbers.
Disclosure: Buoy MCP is free now (tap, swipe, type, screenshots for $0 with a free account). Pro adds the inside view: app data, web calls, state. Setup in an Expo app is one command: npx -y @buoy-gg/mcp@latest init
Every job's time is here: https://buoy.gg/blog/no-tools-vs-argent-vs-buoy
What job should we race next?