[24 SEP 2026][6 MIN READ]

SOTA Without the Token Bill: 95% Success at $0.21 per Task

[Browser agent benchmark series][Part 2 of 4]

By Jayoo Hwang, Vansh Ramani, and Shourya Vir Jain

Browser agents have advanced immensely, but they remain expensive because they consume many tokens and often rely on costly models. Ramain is a general browser agent that stays reliable even when it uses cheaper models.

95%
Task success
$0.21
Cost per task
36%
Lower cost

Observation is the expensive part

Browser agents perceive pages through the DOM, screenshots, or both. The DOM can be efficient on text-heavy pages, but complex automation often requires spatial understanding that the DOM alone cannot provide.

The strongest general agents combine both representations. Every click can then require thousands of input tokens, which bloats the context window and degrades the model's ability to focus.

How Ramain makes a hybrid agent efficient

Ramain keeps the benefits of screenshot and DOM grounding while reducing the context each action needs.

  • Context management compresses previous steps while preserving task-relevant details.
  • A hierarchical page representation gives the agent a compact structured decomposition of large pages.
  • Dynamic state tracking reports precise browser changes caused by each action.

Same model, higher success, lower cost

We evaluated Ramain and the open-source browser agent browser-use on the same difficult enterprise tasks. Both agents used Gemini 3.8 Flash.

Ramain achieved 95% success at $0.21 per task. Browser-use achieved 85% at $0.33. That is a 10-percentage-point improvement in success while costing 36% less per task.

Ramain reaches 95 percent success at 21 cents per task versus browser-use at 85 percent success and 33 cents per task, both using Gemini 3.8 Flash.
On the same model and enterprise task set, Ramain is both more reliable and less expensive than browser-use.

Cost decides which workflows get automated

For an operations team, cost per task determines whether automation clears the business-case threshold. Expensive implementations leave teams automating their three worst workflows instead of all thirty.

A more efficient agent moves that line. Workflows that never justified an RPA quote become practical, while the team preserves accuracy on the long browser tasks that matter.

Key takeaway

Context efficiency lets Ramain use fewer tokens and deliver strong performance with cheaper models.