Title: Models · Steerbay
URL: https://www.steerbay.com/en/models/
Language: English (Česky: https://www.steerbay.comhttps://www.steerbay.com/cs/modely/; x-default: https://www.steerbay.comhttps://www.steerbay.com/cs/modely/)
Summary: Every model on Steerbay, rated: web navigation, long tasks, computer use, office work, speed and safety, what each is best for, which reasoning level to use and what it costs.

---

Models

# Which model for which work.

Every model on Steerbay, rated for the work agents do here. Open one to see why.

Claude Opus 5.5 Anthropic · Pro and Max The best for long, demanding work: research, analysis and tasks that run for hours. Cost Cost: medium, 3 of 5 Overall 9.1 out of 10

Our strongest model. Reach for it when a task is long, open-ended or has to be right first time.

- **Best for**: - Tasks that run for hours - Research and analysis with a finished result
- **Not for**: - Short, routine tasks
- **Reasoning**: High; Max only for the hardest tasks

### Ratings

- **Web navigation**: 9.0
- **Long tasks**: 9.5
- **Computer use**: 9.0
- **Office work**: 9.5
- **Speed**: 7.5
- **Safety**: 9.5

### What the ratings rest on

Claude Opus 5.5: published results and their sources

| Benchmark | Result | Source |
| --- | --- | --- |
| Long tasks |  |  |
| Terminal-Bench 4.0 Long tasks in a terminal, run independently | 65.15%, 1st of 45 | [Vals AI (new tab)](https://www.vals.ai/benchmarks/terminal-bench-4) |
| Arena Agent Tool use in real agent tasks | +14.33%, 1st | [Arena (new tab)](https://arena.ai/leaderboard/agent) |
| Vending-Bench 2 A simulated year running a shop | $9,235 balance, 9th | [Andon Labs (new tab)](https://andonlabs.com/evals/vending-bench-2) |
| Computer use |  |  |
| OSWorld 2.1 108 long tasks on a computer, measured by Anthropic | 81.8% | [Anthropic (new tab)](https://www.anthropic.com/claude-opus-5-5) |
| CUA-bench Control from screen pixels alone, run independently | 14.00%, 2nd of 8 | [Vals AI (new tab)](https://vals.ai/models/anthropic_claude-opus-5-5) |
| Office work |  |  |
| GDPval-AA v2.1 Professional assignments, compared head to head | Elo 1866, 1st | [Artificial Analysis (new tab)](https://artificialanalysis.ai/evaluations/gdpval-aa) |
| Speed |  |  |
| Artificial Analysis Output speed | 96 tokens/s | [Artificial Analysis (new tab)](https://artificialanalysis.ai/leaderboards/models) |
| Safety |  |  |
| Prompt injection Successful attacks by planted instructions in a browser, measured by Anthropic | 0.09% without safeguards, 0% with | [Anthropic (new tab)](https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba5024219938094694bd/Claude%20Opus%205.5%20System%20Card.pdf) |

### Price and facts

- **On plans**: Pro and Max
- **Allowance use**: Uses your allowance twice as fast as Claude Sonnet 5.5.
- **Provider list price**: $4 per million input tokens, $20 per million output tokens
- **Released**: 22 September 2026
- **Context window**: 1 million tokens

Sol 6.1 OpenAI · Plus, Pro and Max Our pick for work on the web: first in the independent browsing test, at a low cost. Cost Cost: low, 2 of 5 Overall 8.8 out of 10

Our recommended model for most web work. It beat every other model in the independent browsing test.

- **Best for**: - Searching and comparing prices on websites - Forms and multi-step processes
- **Not for**: - The longest open-ended tasks
- **Reasoning**: Medium on the web; High for long tasks

### Ratings

- **Web navigation**: 9.5
- **Long tasks**: 8.5
- **Computer use**: 8.5
- **Office work**: 8.0
- **Speed**: 8.5
- **Safety**: 8.0

### What the ratings rest on

Sol 6.1: published results and their sources

| Benchmark | Result | Source |
| --- | --- | --- |
| Web navigation |  |  |
| Browser Use Bench v2.1 180 long tasks on live websites | 79.6, 1st | [Browser Use (new tab)](https://browser-use.com/benchmarks/agents) |
| Long tasks |  |  |
| SRE Bench Resolving server incidents, run independently | 50.76%, 2nd of 33 | [Vals AI (new tab)](https://vals.ai/models/openai_gpt-6.1-sol) |
| Terminal-Bench 4.0 Long tasks in a terminal, run independently | 55.05%, 6th of 45 | [Vals AI (new tab)](https://www.vals.ai/benchmarks/terminal-bench-4) |
| Arena Agent Tool use in real agent tasks | +11.72%, 5th | [Arena (new tab)](https://arena.ai/leaderboard/agent) |
| Computer use |  |  |
| OSWorld 2.0 Tasks on a computer, offline set, measured by OpenAI | 71.4% | [OpenAI (new tab)](https://openai.com/index/introducing-gpt-6-1-sol/) |
| Office work |  |  |
| GDPval-AA v2.1 Professional assignments, compared head to head | Elo 1592, 33rd | [Artificial Analysis (new tab)](https://artificialanalysis.ai/evaluations/gdpval-aa) |
| Speed |  |  |
| Browser Use Bench v2.1 Average time per web task | 8 minutes | [Browser Use (new tab)](https://browser-use.com/benchmarks/agents) |
| Safety |  |  |
| Safety stress test Unsafe actions in a computer-use stress test, measured by OpenAI (lower is better) | 4.3% | [OpenAI (new tab)](https://openai.com/index/introducing-gpt-6-1-sol/) |

### Price and facts

- **On plans**: Plus, Pro and Max
- **Allowance use**: Uses your allowance at the same rate as Claude Sonnet 5.5.
- **Provider list price**: $2 per million input tokens, $10 per million output tokens
- **Full name**: GPT-6.1 Sol
- **Released**: 29 September 2026
- **Context window**: 1.05 million tokens

Claude Sonnet 5.5 Anthropic · Free, Plus, Pro and Max A dependable everyday model, the safest in a browser. It is on every plan. Cost Cost: low, 2 of 5 Overall 8.7 out of 10

A safe choice for everyday tasks, documents and work on unfamiliar websites.

- **Best for**: - Documents, spreadsheets and presentations - Work on unfamiliar websites
- **Not for**: - Open-ended tasks of many hours
- **Reasoning**: High; Extra high for long tasks

### Ratings

- **Web navigation**: 8.0
- **Long tasks**: 9.0
- **Computer use**: 8.5
- **Office work**: 9.5
- **Speed**: 8.5
- **Safety**: 10

### What the ratings rest on

Claude Sonnet 5.5: published results and their sources

| Benchmark | Result | Source |
| --- | --- | --- |
| Long tasks |  |  |
| Terminal-Bench 4.0 Long tasks in a terminal, run independently | 64.14%, 2nd of 45 | [Vals AI (new tab)](https://www.vals.ai/benchmarks/terminal-bench-4) |
| Vending-Bench 2 A simulated year running a shop | $11,165 balance, 5th | [Andon Labs (new tab)](https://andonlabs.com/evals/vending-bench-2) |
| Arena Agent Tool use in real agent tasks | +11.95%, 4th | [Arena (new tab)](https://arena.ai/leaderboard/agent) |
| Computer use |  |  |
| OSWorld 2.1 108 long tasks on a computer, measured by Anthropic | 80.1% | [Anthropic (new tab)](https://www-cdn.anthropic.com/870c8f525702625d2c62fc6dd04c857e3250bec1/Claude%20Sonnet%205.5%20System%20Card.pdf) |
| Office work |  |  |
| GDPval-AA v2.1 Professional assignments, compared head to head | Elo 1838, 2nd | [Artificial Analysis (new tab)](https://artificialanalysis.ai/evaluations/gdpval-aa) |
| Speed |  |  |
| Artificial Analysis Output speed | 141 tokens/s | [Artificial Analysis (new tab)](https://artificialanalysis.ai/leaderboards/models) |
| Safety |  |  |
| Prompt injection Successful attacks by planted instructions in a browser, measured by Anthropic | 0% with or without safeguards | [Anthropic (new tab)](https://www-cdn.anthropic.com/870c8f525702625d2c62fc6dd04c857e3250bec1/Claude%20Sonnet%205.5%20System%20Card.pdf) |

### Price and facts

- **On plans**: Free, Plus, Pro and Max
- **Allowance use**: The reference rate: the other models are compared with it.
- **Provider list price**: $2 per million input tokens, $10 per million output tokens
- **Released**: 28 September 2026
- **Context window**: 1 million tokens

Astra 6 OpenAI · Pro and Max Very capable, but slow and costly for browsing. Keep it for the hardest work on a computer. Cost Cost: very high, 5 of 5 Overall 8.4 out of 10

Strong on a computer and in long operations, but more model than browsing needs. Sol 6.1 is better on the web.

- **Best for**: - The hardest tasks in apps - Work Sol 6.1 could not finish
- **Not for**: - Everyday browsing
- **Reasoning**: Medium on the web; High on a computer

### Ratings

- **Web navigation**: 8.5
- **Long tasks**: 8.5
- **Computer use**: 9.0
- **Office work**: 8.0
- **Speed**: 7.0
- **Safety**: 8.5

### What the ratings rest on

Astra 6: published results and their sources

| Benchmark | Result | Source |
| --- | --- | --- |
| Web navigation |  |  |
| Browser Use Bench v2.1 180 long tasks on live websites | 77.9, 2nd | [Browser Use (new tab)](https://browser-use.com/benchmarks/agents) |
| Long tasks |  |  |
| Vending-Bench 2 A simulated year running a shop | $15,515 balance, 1st | [Andon Labs (new tab)](https://andonlabs.com/evals/vending-bench-2) |
| Terminal-Bench 4.0 Long tasks in a terminal, run independently | 59.60%, 3rd of 45 | [Vals AI (new tab)](https://www.vals.ai/benchmarks/terminal-bench-4) |
| Arena Agent Tool use in real agent tasks | +13.09%, 2nd | [Arena (new tab)](https://arena.ai/leaderboard/agent) |
| Computer use |  |  |
| CUA-bench Control from screen pixels alone, run independently | 19.17%, 1st of 8 | [Vals AI (new tab)](https://vals.ai/models/openai_gpt-6-astra) |
| Office work |  |  |
| GDPval-AA v2.1 Professional assignments, compared head to head | Elo 1574, 37th | [Artificial Analysis (new tab)](https://artificialanalysis.ai/evaluations/gdpval-aa) |
| Speed |  |  |
| Artificial Analysis Output speed | 47 tokens/s | [Artificial Analysis (new tab)](https://artificialanalysis.ai/leaderboards/models) |
| Safety |  |  |
| Safety stress test Unsafe actions in a computer-use stress test, measured by OpenAI (lower is better) | 2.4% | [OpenAI (new tab)](https://openai.com/index/introducing-gpt-6-1-sol/) |

### Price and facts

- **On plans**: Pro and Max
- **Allowance use**: Uses your allowance five times as fast as Claude Sonnet 5.5.
- **Provider list price**: $10 per million input tokens, $50 per million output tokens
- **Full name**: GPT-6 Astra
- **Released**: 3 September 2026
- **Context window**: 1.05 million tokens

Luna 6 OpenAI · Plus, Pro and Max Fast and very cheap for simple jobs: look-ups and sorting, not long work. Cost Cost: very low, 1 of 5 Overall 6.4 out of 10

For simple work in volume. Give anything with many steps to Sol 6.1 or Claude Sonnet 5.5.

- **Best for**: - Quickly looking up one fact - Sorting and short summaries
- **Not for**: - Tasks with many steps
- **Reasoning**: High to Extra high; Low only for one step

### Ratings

- **Web navigation**: 7.0
- **Long tasks**: 5.0
- **Computer use**: 5.5
- **Office work**: 6.5
- **Speed**: 9.0
- **Safety**: 5.5

### What the ratings rest on

Luna 6: published results and their sources

| Benchmark | Result | Source |
| --- | --- | --- |
| Web navigation |  |  |
| Browser Use Bench v2.1 180 long tasks on live websites | 69.0, 12th | [Browser Use (new tab)](https://browser-use.com/benchmarks/agents) |
| Long tasks |  |  |
| Terminal-Bench 4.0 Long tasks in a terminal, run independently | 13.64%, 29th of 45 | [Vals AI (new tab)](https://www.vals.ai/benchmarks/terminal-bench-4) |
| Arena Agent Tool use in real agent tasks | +1.93%, 25th | [Arena (new tab)](https://arena.ai/leaderboard/agent) |
| Computer use |  |  |
| OSWorld 2.0 Tasks on a computer, Max reasoning, measured by OpenAI | 52.7% | [OpenAI (new tab)](https://openai.com/index/introducing-gpt-6-sol-and-luna/) |
| Office work |  |  |
| GDPval-AA v2.1 Professional assignments, compared head to head | Elo 1432, 64th | [Artificial Analysis (new tab)](https://artificialanalysis.ai/evaluations/gdpval-aa) |
| Speed |  |  |
| Artificial Analysis Output speed | 139 tokens/s | [Artificial Analysis (new tab)](https://artificialanalysis.ai/leaderboards/models) |
| Safety |  |  |
| Safety stress test Unsafe actions in a computer-use stress test, measured by OpenAI (lower is better) | 13.7% | [OpenAI (new tab)](https://openai.com/index/introducing-gpt-6-1-sol/) |

### Price and facts

- **On plans**: Plus, Pro and Max
- **Allowance use**: Uses your allowance about twenty times more slowly than Claude Sonnet 5.5.
- **Provider list price**: $0.10 per million input tokens, $0.50 per million output tokens
- **Full name**: GPT-6 Luna
- **Released**: 22 September 2026
- **Context window**: 1.05 million tokens

- excellent 9+
- strong 8+
- fair 7+
- weak below 7

Marks out of 10. Last reviewed 11 October 2026.

Comparison

## Every model, side by side.

Model ratings by area, reasoning, cost and plans

| Model | Overall | Web navigation | Long tasks | Computer use | Office work | Speed | Safety | Reasoning | Cost | Plans |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| [Claude Opus 5.5](https://www.steerbay.com/en/models/#opus-5-5) | 9.1 | 9.0 | 9.5 | 9.0 | 9.5 | 7.5 | 9.5 | High; Max only for the hardest tasks | Cost: medium, 3 of 5 | Pro and Max |
| [Sol 6.1](https://www.steerbay.com/en/models/#sol-6-1) | 8.8 | 9.5 | 8.5 | 8.5 | 8.0 | 8.5 | 8.0 | Medium on the web; High for long tasks | Cost: low, 2 of 5 | Plus, Pro and Max |
| [Claude Sonnet 5.5](https://www.steerbay.com/en/models/#sonnet-5-5) | 8.7 | 8.0 | 9.0 | 8.5 | 9.5 | 8.5 | 10 | High; Extra high for long tasks | Cost: low, 2 of 5 | Free, Plus, Pro and Max |
| [Astra 6](https://www.steerbay.com/en/models/#astra-6) | 8.4 | 8.5 | 8.5 | 9.0 | 8.0 | 7.0 | 8.5 | Medium on the web; High on a computer | Cost: very high, 5 of 5 | Pro and Max |
| [Luna 6](https://www.steerbay.com/en/models/#luna-6) | 6.4 | 7.0 | 5.0 | 5.5 | 6.5 | 9.0 | 5.5 | High to Extra high; Low only for one step | Cost: very low, 1 of 5 | Plus, Pro and Max |

How we rate

## Six things that matter.

The overall rating is a weighted average of six marks. The web counts most, because that is where agents work most. Price is not part of a rating.

- **35% Web navigation**: Finds and does what it should on real websites.
- **20% Long tasks**: Stays on work of dozens of steps and finishes it.
- **15% Computer use**: Operates apps from what it sees on the screen.
- **10% Office work**: Documents, spreadsheets, research and analysis.
- **10% Speed**: How quickly it answers and finishes a task.
- **10% Safety**: Ignores instructions planted on websites.

The full reasons behind each mark and advice on reasoning are in the guide [Models](https://www.steerbay.com/en/docs/models/); plans are in [Pricing](https://www.steerbay.com/en/pricing/).

## Try them on your own work.

Change the model at any time; your files, conversations and sign-ins stay. Not sure? Ask the AI help.

[Open the app](https://app.steerbay.com/)
