AI Agents

What an AI agent turn really costs

Price per million tokens is the number everyone compares, and on its own it tells you almost nothing about your bill. What matters is how many tokens your agent sends per turn, times how many turns you take. We measured the first part on our own servers, and the answer surprised us: most of every turn is the agent talking to itself.

The measurement

Measured input tokens per agent turn
AgentVersionInput tokens per turnHowMeasured
Hermes Agent0.21.xabout 15,000one tool-using turn, Sonnet 4.617 Sep 2026
OpenClaw2026.9.2about 24,500openclaw agent --message, clean home directory19 Sep 2026

Nearly all of those tokens are the agent’s fixed system prompt and tool definitions, re-sent on every turn. Your actual message is a rounding error on top. So cost barely depends on what you asked: a “hello” costs nearly as much as a refactor. OpenClaw’s prompt is about 1.6× Hermes’s.

What we actually paid: about $0.05 for one Hermes turn on Sonnet 4.6, and $0.115 for a three-call tool run.

Worked examples

The arithmetic below is ours, on the measured token counts, at a list price of $3 per million input tokens (the vendor’s published Sonnet input rate when we measured), with no prompt caching and output tokens left out. Check your provider’s current page before relying on the rate.

Hermes Agent on Sonnet

15,000 tokens x $3 / 1,000,000   = $0.045 per turn   (we observed about $0.05)
30 turns a day                    = about $1.35 a day
30 days                           = about $40 a month, before output tokens

OpenClaw on Sonnet

24,500 tokens x $3 / 1,000,000   = $0.0735, about $0.07 per turn
30 turns a day                    = about $2.20 a day
30 days                           = about $66 a month, before output tokens

Either agent on a cheaper model

The token count does not change when you change model; only the rate does. A model at a tenth of Sonnet’s input price makes the same Hermes turn about $0.0045 and the same OpenClaw turn about $0.007. That is the whole case for tiering: put the fixed prompt on a cheap model for routine work and pay Sonnet rates only where Sonnet earns them. We do this ourselves. Every validation run that is not specifically testing Sonnet uses gpt-4o-mini.

For scale against the server: at $0.05 a turn, a $5-a-month VPS costs the same as about 100 turns. The server is not the bill to optimise; see the cheapest VPS for AI agents.

Embeddings: nearly free, easy to misroute

Agent memory turns your files into embeddings so it can search them. On 29 September 2026 (run 20260929-215021) we indexed two files with OpenClaw’s memory plugin and ran a search that returned the right line at a score of 0.607. The whole thing used 48 tokens of text-embedding-3-small: about $0.00000096 at list price, $0.0000012 through our gateway with its 20% surcharge.

The price is not the problem. The routing is: by default the plugin sent the key to api.openai.com, which rejected it with 401, because the key belonged to a different OpenAI-compatible endpoint. Chat worked; memory silently did not. The fix is in self-host OpenClaw. The question to ask of any agent’s memory is whether it works with a non-OpenAI base URL, not what it costs.

Cheap-looking failures that were really billing

  • The provider reserves your maximum, not your usage. OpenRouter reserves the request’s full max_tokens up front, and Hermes asks for 65,536 on Sonnet 4.6. On 16 September 2026, with $6.92 of $135 left, those calls got 402 “requested up to 65536 tokens, but can only afford 6849” while short calls passed and health checks stayed green.
  • The agent may report it as something else. Hermes showed that 402 as Context length exceeded (43 tokens). Cannot compress further. Check the provider’s HTTP status before debugging the context window.
  • Low balance stops the expensive model first. With $0.095 left under a key’s own limit (30 September 2026), cheap-model completions in our canary still passed. An agent that “still works” can be one that has quietly lost its best model.
  • Free models churn. On 27 September 2026 minimax/minimax-m3:free began returning 404: OpenRouter had retired the free variant and nothing warned us. We replaced it with the paid row by hand. If you lean on :free models, check they are still served and keep a paid fallback.

What these numbers are not

  • Not one controlled experiment. The Hermes and OpenClaw counts come from different dates, with tools enabled, one to three runs each, no variance.
  • Not every mode. A Hermes one-shot (-z) on 30 August 2026 logged 4,326 input tokens. That is a different mode, and we have not reconciled it with the 15,000.
  • Not cached. Dollar figures assume list price with no prompt caching. We did not measure caching.
  • Not Claude Code or Codex. We have not measured their per-turn tokens, so they are not on this page.
  • Not a quality judgement. Our agent test chats used gpt-4o-mini; nothing here says which model writes better code.

Run either agent with model choice built in

AgentOcean installs Hermes and OpenClaw pre-wired to an OpenAI-compatible gateway, so switching the model behind the fixed prompt is a setting, not a migration.

Related

Measured on

  • Hermes Agent 0.21.x turn: 17 September 2026, Sonnet 4.6.
  • OpenClaw 2026.9.2 turn: 19 September 2026.
  • OpenRouter 402 behind “context length exceeded”: 16 September 2026.
  • Free model retired: 27 September 2026.
  • Embeddings: run 20260929-215021, 29 September 2026, 386 s, $0.04 for the whole server run.
  • Cheap completions passing under a near-empty key: canary-20260930-003625, 30 September 2026.

On servers and accounts we paid for. Token counts are single observations; derived dollar figures are labelled as arithmetic.