Budget

Token and cost calculator

Reflection has not published API prices for the beta, so enter your own. The calculator shows where tokens go, especially reasoning, which Beam always produces.

API usage
Beam's reasoning length depends on the task and effort. Read usage.completion_tokens_details.reasoning_tokens from your responses and enter the average.

Leave empty to see tokens only. Reasoning is billed as output on OpenAI-style APIs; check Reflection's terms when prices appear.

API usage
Tokens per request
4,100
Tokens per day
4.1M
Tokens per month
123M
Average tokens per minute
2,847
Reasoning share of output
71%
API cost per month
enter prices
Input tokens per requestAnswer tokens per requestReasoning tokens per request
Self-hosting
Total output tokens per second across all requests on the server. Measure it with your engine and batch size; it varies widely.
Self-hosting
Cost per month
$17,280
Output tokens per month
3.9B
Cost per 1M output tokens
$4.44
Rough token counter

About 0 tokens · 0 characters

Beam's tokenizer is not public yet. This uses about 4 characters per token for Latin text and about 1 token per character for Chinese, Japanese and Korean.

Rate limits in short
  • Limits apply to the whole organization, shared by all projects and keys.
  • They count requests and tokens per minute and per UTC day; reasoning tokens count too.
  • Concurrent requests are capped per organization and per key.
  • 429 with code rate_limit_exceeded: wait at least Retry-After seconds. x-should-retry: false means the daily budget is spent until 00:00 UTC.
  • Watch x-ratelimit-remaining-requests and x-ratelimit-remaining-tokens (and their -day versions) to slow down early.

Source: developers.reflection.ai/rate-limits