> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.empiriolabs.ai/account-usage-api/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.empiriolabs.ai/_mcp/server. # Account Usage API Use account endpoints when you want to reconcile spend, show usage inside your own product, or export saved Playground tests. ## Usage and balance `GET /v1/account/usage` returns your current credit balance, a usage summary for the query window, an account-level spend breakdown by product, and recent usage events for the account attached to the API key. ```bash curl "https://api.empiriolabs.ai/v1/account/usage?limit=25" \ -H "Authorization: Bearer $EMPIRIOLABS_API_KEY" ``` ### Query parameters | Parameter | Meaning | | ------------- | ------------------------------------------------------------------------------------------------------------------ | | `limit` | Events per page. Default 100, maximum 200. | | `before` | Return events created before this timestamp. Pass the previous response's `next_cursor` here to page further back. | | `from` / `to` | Time window for events. The `spend` block honors the same window. | | `model` | Filter to one model ID. Also useful for a per-model `cache_hit_rate`. | | `status` | `success` or `error`. | | `source` | `api`, `playground`, `compose`, `gpu_cloud`, or `hosted_agents`. | ### Response shape ```json { "object": "account_usage", "balance": { "amount": 42.5, "currency": "USD", "auto_topup_enabled": true, "auto_topup_threshold": 10, "auto_topup_amount": 50, "updated_at": "2026-07-10T18:04:11Z" }, "plan": { "slug": "pro", "display_name": "Pro", "status": "active", "billing_cadence": "monthly", "current_period_end": "2026-09-08T18:04:11Z", "cancel_at_period_end": false, "window": { "start": "2026-08-08T18:04:11Z", "end": "2026-08-15T18:04:11Z" } }, "summary": { "requests": 25, "total_cost": 0.8412, "total_tokens": 154200, "input_tokens": 120400, "output_tokens": 33800, "cache_read_tokens": 41200, "cache_write_tokens": 0, "cache_hit_rate": 0.3422, "errors": 1, "avg_latency_ms": 2140 }, "spend": { "currency": "USD", "total": 18.4029, "model": 12.1103, "compose": 1.2926, "gpu_cloud": 4.2, "hosted_agents": 0.8, "from": null, "to": null }, "data": [ ... ], "has_more": true, "next_cursor": "2026-07-08T19:22:41.512Z" } ``` | Field | Meaning | | -------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `balance` | Current credit balance in USD, plus your auto top-up settings (`auto_topup_enabled`, `auto_topup_threshold`, `auto_topup_amount`). | | `plan` | Current subscription tier, status, billing cadence, paid-through date, cancellation state, and weekly window. Present for an account with a current active or past-due subscription. A past-due plan grants no plan access. Allowance limits remain on the Billing and pricing pages rather than in this API response. | | `summary` | Totals across the whole query window: request count, cost, token counters, `cache_hit_rate`, error count, and average latency. Honors the same `from`/`to`/`model`/`status`/`source` filters as the events (but not `before`, since the window total is stable across pages), so its counts are true totals, not the size of the returned page. | | `spend` | Account-level spend split by product from the billing ledger: `model` (model and API usage, including Playground), `compose`, `gpu_cloud`, and `hosted_agents`. Honors `from`/`to` when set, otherwise all-time. Matches the dashboard Usage page. | | `data` | Usage events, newest first. | | `has_more` / `next_cursor` | Pagination. When `has_more` is true, pass `next_cursor` as the `before` parameter of the next request. | `summary.cache_hit_rate` is the share of input tokens served from prompt cache across the query window (`cache_read / input`, 0 to 1, `null` when the window has no input tokens). Cached input bills at the model's cache-read rate, so this is your effective input-cost discount signal. Filter by `model` to get a per-model rolling rate. ### Usage events Each usage event includes: | Field | Meaning | | ----------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `source` | `api`, `playground`, `compose`, `gpu_cloud`, `hosted_agents`, or `template` (a generation that used an effect template; template rows carry the template slug under `metadata.template`). The `source` query filter accepts the first five values. | | `model` and `endpoint` | What was called. For GPU Cloud rows, `model` is the GPU SKU identifier. | | `tokens` | Input, output, cached, and total token counts. Non-token-billed rows, such as GPU Cloud runtime events, return zero token counts and put runtime details in `metadata`. | | `cost.amount` | Debited pay-as-you-go credits in USD. This is the single source of truth for what was charged to the credit balance. | | `cost.included` / `cost.partial` / `cost.plan` | `included: true` means the plan covered the whole request and `amount` is zero. `partial: true` means the plan covered part and `amount` is the pay-as-you-go remainder. `plan` is the tier on the account when that request settled, so historical rows do not change after an upgrade or downgrade. | | `status`, `status_code`, `error` | Request outcome | | `output_urls` | Generated media URLs when present | | `tool_breakdown` | Array of `{name, count, unit_cost_usd, subtotal_usd}` entries when the model billed per-tool surcharges and the catalog knows the per-tool rate. `null` when no tools fired. Same per-tool subtotals the dashboard's expanded usage row shows. | | `metadata.compose` | Present on Compose parent rows. Includes `recipe`, `mode`, `phase`, `total_cost_usd`, `leg_count`, and `legs` for the model stages that made up the production. | | `metadata.agent_type` / `metadata.agent_name` / `metadata.agent_instance_id` | Present on Hosted Agent model calls so you can trace usage back to the agent. | | `metadata.gpu_slug` / `metadata.gpu_display` / `metadata.num_gpus` / `metadata.seconds` / `metadata.price_hourly` | Present on GPU Cloud runtime events so you can reconcile the display name, GPU count, runtime, and hourly rate. | | `metadata.tool_usage` | Map of `{tool_name: count}` when a model billed per-tool surcharges (e.g. `web_search`, `image_search`, `code_interpreter`, `web_extractor`). Only present when at least one tool fired. | | `metadata.worker_usage` | The raw usage object the model service reported, including tier labels such as `pricing_tier_label` and per-dimension counters (e.g. `citation_tokens`, `num_search_queries`). Present when the model reported extended usage. | | `metadata.list_cost` / `metadata.billed_cost` / `metadata.discount_amount` | Only present when an account-level pricing discount was applied. `cost.amount` already reflects the post-discount amount; these fields document the discount for receipts. | `cost.amount` is always the final debited amount. The `metadata.list_cost` / `metadata.billed_cost` pair is informational. When a discount is applied, `cost.amount` equals `billed_cost`. When no discount applies, neither field is set; just read `cost.amount`. For subscription usage, `cost.included` and `cost.partial` are mutually exclusive. A partly covered request starts only when the available credit balance can cover its full pay-as-you-go remainder. If not, it returns `402 insufficient_credits` and creates no usage charge or plan draw. For GPU Cloud events, use `source: "gpu_cloud"` plus `metadata.seconds` and `metadata.price_hourly` for runtime reporting. The `tokens` object remains in the response for schema stability, but GPU runtime rows are not token-billed. For Compose events, use `source: "compose"` to retrieve full-video production rows: ```bash curl "https://api.empiriolabs.ai/v1/account/usage?source=compose&limit=20" \ -H "Authorization: Bearer $EMPIRIOLABS_API_KEY" ``` A Compose row uses `endpoint: "/v1/videos/compose"` and stores the production breakdown under `metadata.compose`: ```json { "object": "usage_event", "source": "compose", "model": "custom", "endpoint": "/v1/videos/compose", "cost": { "amount": 0.1832, "currency": "USD" }, "metadata": { "category": "compose", "compose": { "recipe": "custom", "mode": "render", "phase": "render", "total_cost_usd": 0.1832, "leg_count": 5, "legs": [ { "stage": "voice", "model": "inworld-tts-mini", "cost_usd": 0.012 }, { "stage": "visual", "model": "seedance-2-0-pro", "cost_usd": 0.14 } ] } } } ``` Read `cost.amount` for the final debited amount. Use `metadata.compose.legs` when you want to show which model stages contributed. ### Tool usage example When you call a model that bills per-tool surcharges (Qwen, Perplexity, MiMo, Mistral) and the model invokes those tools, the response includes a normalized `tool_usage` map: ```json { "id": "...", "object": "usage_event", "model": "qwen3-6-plus", "cost": { "amount": 0.084, "currency": "USD" }, "tokens": { "input": 1250, "output": 480, "total": 1730 }, "tool_breakdown": [ { "name": "web_search", "count": 2, "unit_cost_usd": 0.026, "subtotal_usd": 0.052 }, { "name": "image_search", "count": 1, "unit_cost_usd": 0.0208, "subtotal_usd": 0.0208 } ], "metadata": { "tool_usage": { "web_search": 2, "image_search": 1 } } } ``` The `cost.amount` already includes the per-tool surcharges (here: `2 × $0.026 web_search + 1 × $0.0208 image_search + token cost`). `metadata.tool_usage` is the raw invocation-count map; `tool_breakdown` adds the per-tool unit cost and subtotal so you can render a billing breakdown without fetching catalog pricing yourself. ## Saved Playground chats The public API exposes saved Playground conversations as read-only resources. List your saved Playground conversations: ```bash curl "https://api.empiriolabs.ai/v1/playground/conversations" \ -H "Authorization: Bearer $EMPIRIOLABS_API_KEY" ``` Retrieve a single conversation by ID, including its full message history: ```bash curl "https://api.empiriolabs.ai/v1/playground/conversations/CONVERSATION_ID" \ -H "Authorization: Bearer $EMPIRIOLABS_API_KEY" ``` Saving and deleting Playground chats still happens in the [dashboard Playground](https://platform.empiriolabs.ai/dashboard/playground). > Query balance, request history, usage counters, costs, and saved Playground chats