By Jayoo Hwang, Vansh Ramani, and Shourya Vir Jain
Cheaper models are not smarter than frontier models. For most long browser tasks, however, endurance matters more than raw intelligence.
- 89.4%
- Highest success rate
- $0.16
- Lowest task cost
- 37K
- Context tokens per task
Long browser tasks are mostly repetition
A 100-step enterprise workflow often looks like this: log into a portal, open the next claim, copy three fields into a tracker, set a status, and repeat forty times. Each individual step is easy. The task is long because it repeats.
Our intuition about long-horizon work comes from coding and mathematics, where later steps can require deep global understanding. In browser work, step forty is usually another form. Duration compounds, not depth.
Context rot, not model size, breaks the run
Each step appends observations. The context fills with stale pages and dead history, and the model begins attending to the wrong information. It repeats finished steps or loses track of the current one.
Paying for a frontier model does not remove this failure mode. A larger context window can simply give the noise more room.
Ramain keeps late steps as reliable as early steps
Ramain keeps summaries compact, updates observations incrementally, and compresses history to what still matters. Past a certain horizon, context management matters more than model size, and context management is a harness-design problem.
Every point on the frontier is a Ramain configuration
We tested Ramain, Claude Code, and Codex on the same long-horizon enterprise browser tasks. On GPT-6 Astra, Ramain reached 89.4% success at $1.04 per task, against Codex at 88.2% and $2.55. On Fable 5.1, Ramain reached 88.2% at $0.79, against Claude Code at 87.1% and $1.38.
The cheapest configuration shows how far the harness carries a small model. Ramain on Gemini 3.8 Flash reached 85.9% success at $0.16 per task, within 1.2 points of Claude Code on Fable 5.1 at nearly nine times lower cost, and within 2.3 points of Codex on GPT-6 Astra at sixteen times lower cost.
With the same Gemini 3.8 Flash model, Ramain used 37,000 context tokens per task on average versus 165,000 for Codex. Ramain achieved higher performance while costing four times less.

Ramain's context-efficient architecture sits on the cost-performance frontier for long, tedious browser tasks.