Beta coming soon. Join the release list below.

Rollout playbook

Tune Routing Policy with Evidence

Start with a high-volume workload in monitor mode. Compare cost, capability, cacheability, and evaluation evidence before allowing FrugalAI to route an eligible request.

Discuss a pilot workload

Served catalog coverage

Nine providers in the served catalog

OpenAI
Anthropic
Google
DeepSeek
Mistral
Groq
Together
Fireworks
xAI

The catalog includes 31 chat models. A production route only considers providers configured with your organization's keys.

Controlled rollout

Quality stays in the loop

01

Monitor

Observe unchanged requests

02

Assist

Route unpinned requests

03

Enforce

Apply approved policy

Sampled evaluation runs inside assist and enforce; a configured SLO breach returns the gateway to monitor mode.

Tune routing policy with evidence

Start with a high-volume workload in monitor mode, compare eligible alternatives against capability and evaluation evidence, then advance only the routes that satisfy the policy you approved.

Product and access

Frequently asked questions

Direct answers about routing, evidence, billing, and the current owner-led beta.

Ask about a pilot

gpt-4.1

claude-fable-5

auto

Decision receipt

gemini-3-flash

Reason
lowest eligible
Mode
assist
Fallback
ready

Cost-aware model routing

Select the lowest-cost model that remains eligible under capability, policy, provider, and evidence checks.

See rollout details

Controlled rollout

Quality stays in the loop

01

Monitor

Observe unchanged requests

02

Assist

Route unpinned requests

03

Enforce

Apply approved policy

Sampled evaluation runs inside assist and enforce; a configured SLO breach returns the gateway to monitor mode.

Controlled quality rollout

Move from monitor to assist and enforce, with sampled evaluation and SLO rollback as the guardrail.

See rollout details

Bundled demo fixture

Request decision ledger

WorkloadRequestedRoutedReasonCost
support-botclaude-fable-5claude-haiku-4-5enforce / chat$0.002299
codegen-ciautomoonshotai/kimi-k2-instruct-0905exact cache hit$0.000000
rag-searchautogemini-3-flash-previewlowest eligible$0.003301

Request decision ledger

Trace requested and routed models, reasons, cache outcomes, latency, cost, and attribution.

See rollout details

Budget pressure

Act before the invoice

Support58%
Extraction76%
Experiments92%
AlertDowngradeBlock

Team budget pressure

Apply alert, downgrade, or block behavior at explicit organization and team thresholds.

See rollout details

Incoming request

tenant + policy scope

Exact lookup

hit: 23 ms

Safe reuse

provider call avoided

Optional semantic matching is policy-scoped. Sensitive or dynamic requests can bypass cache entirely.

Scope-aware caching

Avoid repeat provider calls with exact cache and optional policy-approved semantic matching.

See rollout details