AnchorShell Relay

Control AI spend
before it gets out of hand.

AnchorShell Relay helps you spend less on AI by prioritizing lower-cost model paths, controlling requests, tokens, and spend, and showing exactly what every user and agent consumed.

Runaway-agent protection

Global capacityavailable
Coding agentdaily spend reached
blocked
Support agentwithin budget
running

Set the limit before the spend happens.

Control requests, tokens, or spend globally or for one user, agent, model, or provider. Apply the limit per second, minute, hour, day, or month.

  • Reserve expensive models for the work that needs them.
  • Give trusted workflows more room.
  • Stop one agent without stopping everyone else.
Read the limits guide →

Set a limit

1 What to controlRequestsTokensSpend
2 How oftenSecondMinuteHourDayMonth
3 Who or whatEveryoneUser or agentModelProvider

Policy summary

TargetCoding agent
ModelGPT-4.1
Limit$5 per day
$5.00 usedNext request blocked

Check the cost before the request runs.

Relay estimates incoming token usage before dispatch and checks it against configured token and spend limits.

Block new requests before their estimate exceeds the remaining allowance. Provider-reported usage can reconcile the estimate after completion when it is available.

Estimated token check

Remaining daily allowance18,000 tokens
Estimated request4,200 tokens

within limit

Remaining2,000 tokens
Estimated request4,200 tokens

blocked before dispatch

Block new requests before their estimate exceeds the remaining allowance.

Manage AI usage across the whole team.

AnchorShell Relay — Managed combines invitations, role presets, custom product permissions, request attribution, and per-user limits in one hosted operating model.

Managers can see organization usage while individual users can be limited to their own usage, logs, queue, and limits.

Team

AnchorShell Relay — Managed
Managed

Team Members

3 active

User controls

Dave · Developer
Active
Requests
100 / day
Tokens
100,000 / day
Spend
$5.00 / day
Model-specific limits
3 configured
Usage visibility
Own usage

Access detail

Usage access
Own
Limit access
View
Provider access
View
Role presets provide a starting point. Managed administrators can assign custom Relay permissions and user-specific request, token, spend, model, or provider limits.

From queue to model to completed request.

Relay waits for configured capacity, skips paths that are cooling or outside their limits, and records the selected route with owner, token, cost, and timing data.

Manually rank local or free OpenAI-compatible paths first and keep an eligible cloud fallback ready.

Realtime request flow

One request moving through Relay
Flow active

Request rq_214 waits in the queue, skips a cooling free endpoint, moves through the available local model, and completes with Maya as owner, 1,284 tokens, and four cents in cost.

Queue

rq_214Maya · waiting
rq_215Research · waiting
rq_216Coding agent · waiting

Available models

Local modellocal/worker
available
Free endpointfree/general
cooling
Cloud providercloud/standard
available

Cooling path skipped

Completed requests

rq_214completed
Owner
Maya
Tokens
1,284
Cost
$0.04

Local model · 1.3s total

Read the queue and pacing guide →

See usage before the invoice arrives.

See requests, tokens, and spend over time. Managed organization visibility can filter attributed usage by user alongside model, provider, and date range.

Start with the trend, narrow the scope, then inspect the exact routed request behind a spike, retry, or failure.

Usage over time

Requests, tokens, and spend across Relay
Requests usage over the selected dayA line chart with one main peak, a smaller later rise, and the latest data point selected.9006003000Latest86 requests9 AM12 PM3 PM6 PMNow

See exactly what your agent sent.

Operational metadata is recorded independently of payload capture. When deeper debugging is needed, an administrator with settings access can enable the global body-storage setting for new traffic.

Payload capture is optional and off by default. Relay does not require raw prompts and responses to retain request ownership, route, status, token, cost, and timing data.

Request history

Operational metadata for routed model calls
Request history with rq_7F2B selected for inspection
RequestOwnerRouteResultTokensTimingCost
rq_7F2ACoding agentlocal/workerBlocked42 ms
rq_7F2BMayacloud/gpt-4.1Completed2,3181.3 s$0.04
rq_7F2CResearchfree/generalCompleted860920 ms$0.00

Store request and response bodies

Global Relay setting · off by default
Metadata only

New request and response bodies are not retained while capture is off.

Fewer failed requests. Less recovery code.

When enabled, Relay can move eligible requests after connection failures, timeouts, or upstream server failures. Throttled paths can cool while queued work waits for an available option.

Provider fallback

Configured model path
Recovered
Requested routegeneral-purpose
first available path
  1. Primary modelcloud/primary
    Provider statusUnhealthy
    skipped
  2. Regional backupcloud/secondary
    Provider statusAvailable
    selected
  3. Local modellocal/worker
    Provider statusAvailable
    ready
  4. Reserve modelcloud/reserve
    Provider statusAvailable
    ready
Relay skips the unhealthy path and continues through the next available configured model.

Self-host Relay or let AnchorShell manage it.

Use AnchorShell Relay — Managed for organization-aware access and user attribution, or run AnchorShell Relay — Self-Hosted as a single-node service with SQLite-backed configuration and an embedded admin UI.