Open the "Project Phoenix" room and send the message "status update".
success1.0
The final screenshot shows the Project Phoenix room with Alice's sent message 'status update' visible with timestamp/checkmark and an empty composer.

Created by
Compare models, tools, cost, and test runs across real iOS app-control tasks.
Completion assigns 1 to success, 0.5 to partial, and 0 to failure. Ties are broken by mean time on successful runs.
| Rank | Configuration | Tool | Completion | Time / run | Cost / run | Task outcomes |
|---|---|---|---|---|---|---|
| 01 | haiku-4.5 (high)· argent v0.15.0 | argentv0.15.0 | 98% | 1m 07s | $0.22 | Passed 95%, Partial 5%, Failed 0% Time / run1m 07sCost / run$0.22 |
| 02 | gpt-5.4-mini (high)· argent v0.15.0 | argentv0.15.0 | 95% | 1m 14s | $0.13 | Passed 92%, Partial 7%, Failed 2% Time / run1m 14sCost / run$0.13 |
| 03 | haiku-4.5 (low)· argent v0.15.0 | argentv0.15.0 | 94% | 58s | $0.20 | Passed 92%, Partial 5%, Failed 3% Time / run58sCost / run$0.20 |
| 04 | gpt-5.4-mini (high)· agent-device v0.17.6 | agent-devicev0.17.6 | 85% | 1m 29s | $0.10 | Passed 77%, Partial 17%, Failed 7% Time / run1m 29sCost / run$0.10 |
| 05 | haiku-4.5 (low)· agent-device v0.17.6 | agent-devicev0.17.6 | 84% | 1m 53s | $0.13 | Passed 80%, Partial 8%, Failed 12% Time / run1m 53sCost / run$0.13 |
| 06 | haiku-4.5 (high)· agent-device v0.17.6 | agent-devicev0.17.6 | 83% | 1m 25s | $0.15 | Passed 75%, Partial 17%, Failed 8% Time / run1m 25sCost / run$0.15 |
| 07 | gpt-5.4-mini (low)· agent-device v0.17.6 | agent-devicev0.17.6 | 82% | 55s | $0.06 | Passed 73%, Partial 18%, Failed 8% Time / run55sCost / run$0.06 |
| 08 | gpt-5.4-mini (low)· argent v0.15.0 | argentv0.15.0 | 82% | 1m 17s | $0.14 | Passed 78%, Partial 8%, Failed 13% Time / run1m 17sCost / run$0.14 |
| 09 | gpt-5.4-mini (high)· no tool | no tool | 46% | 8m 32s | $0.62 | Passed 42%, Partial 8%, Failed 50% Time / run8m 32sCost / run$0.62 |
| 10 | gpt-5.4-mini (low)· no tool | no tool | 16% | 5m 13s | $0.10 | Passed 8%, Partial 15%, Failed 77% Time / run5m 13sCost / run$0.10 |
Completion score by model and tool.
| model | argentv0.15.0 | agent-devicev0.17.6 | no tool |
|---|---|---|---|
| gpt-5.4-mini | 89%n=120 | 84%n=120 | 31%n=120 |
| low | 82%n=60 | 82%n=60 | 16%n=60 |
| high | 95%n=60 | 85%n=60 | 46%n=60 |
| haiku-4.5 | 96%n=120 | 84%n=120 | 13%n=120 |
| low | 94%n=60 | 84%n=60 | 15%n=60 |
| high | 98%n=60 | 83%n=60 | 12%n=60 |
| overall | 92%n=240 | 84%n=240 | 22%n=240 |
Time uses successful runs only.
Average price per task vs completion score. Colour identifies the model, shape the tool; higher and cheaper is better.
Open the "Project Phoenix" room and send the message "status update".
The final screenshot shows the Project Phoenix room with Alice's sent message 'status update' visible with timestamp/checkmark and an empty composer.

The judge is never told which model produced a screenshot, though the action names reveal which tool drove the device.