Zum Inhalt springen

Costs and budgets

Dieser Inhalt ist noch nicht in deiner Sprache verfügbar.

Each step records tokens in and out and its cost. Amounts are tracked as integer micro-USD, so there are no rounding surprises.

Prices come from a configurable table:

[
{ "provider": "openai", "model": "gpt-4.1-mini", "inputPerMTok": 0.4, "outputPerMTok": 1.6 },
{ "provider": "bedrock", "model": "anthropic.*", "inputPerMTok": 3, "outputPerMTok": 15 },
{ "provider": "mcp", "model": "*", "inputPerMTok": 0, "outputPerMTok": 0, "perToolCallUsd": 0.001 }
]

The values above are examples; use your contract prices. Exact model names win over globs. Prices are looked up by provider name first, then by provider kind. simulated/* and ollama/* cost nothing. A call without a matching price is recorded with priced: false, so gaps are visible.

budget:
maxTokens: 50000
maxCostUsd: 0.5
maxSteps: 12
maxToolCalls: 6
timeoutSeconds: 300

Budgets can be set for the pipeline and for each agent; the stricter value wins. When a budget is used up, the control agent stops the run.

Available in 0.1 Monthly team budgets and the cost dashboard. Next release (0.2) Monthly tenant and use-case budgets with hard stop, alerts at 50/80/100 % as events and audit entries. Planned for 0.2 Monthly per-agent budgets and alert delivery to chat and mail.

Available in 0.1 Every cost line carries tenant, agent, use case, run, step, model and provider. Totals can be aggregated by any of these and exported as CSV or JSON; a Prometheus metric with bounded labels covers dashboards. Hard stops exist per run (each agent’s budget), per team and month (0.1) and per tenant and use case and month (next release); monthly per-agent budgets are planned for 0.2. Prices come from the pinned catalog described in your keys, your models.