Route by evidence
Filter by capability and policy, then choose the lowest-cost qualified route. Pin sensitive workloads and keep the reason on every response.
The FinOps-first AI gateway
One OpenAI-compatible endpoint routes each eligible request to the lowest-cost model that clears your quality bar, serves approved repeats from cache, and applies team budgets before spend becomes an invoice surprise.
Monitor first. Compare against your own traffic. Enforce only after the quality evidence holds.
Control the request before it becomes spend
Most AI cost dashboards explain the bill after the fact. FrugalAI makes the budget state, workload requirements, model catalog, cache policy, and quality floor part of the request decision.
Filter by capability and policy, then choose the lowest-cost qualified route. Pin sensitive workloads and keep the reason on every response.
Partition exact and semantic caches by tenant, route, policy, and prompt version. Sensitive or dynamic routes bypass response reuse.
Allocate spend by team and apply progressive pressure to eligible work before a limit turns into customer-facing downtime.
x-frugalai-requested-model: frontier-model
x-frugalai-routed-model: small-qualified-model
x-frugalai-cache: miss
x-frugalai-cost: $0.0000078
x-frugalai-saved: $0.0009422
x-frugalai-decision: classification route cleared small-tier quality floorA controlled rollout
Pass requests through unchanged. Build the allocation and counterfactual savings ledger.
Compare sampled alternative routes against task-specific checks and quality thresholds.
Route only explicit auto-model requests while direct model choices remain pinned.
Apply approved policy by team and workload, with automatic rollback to monitor.
Flat control-plane pricing
Provider usage is billed by your providers. FrugalAI pricing covers the control plane.
Monitor mode, up to 100,000 requests per month, and a savings report.
Start monitor modeRouting, cache, budgets, dashboard, two seats, and up to 5 million requests per month.
SSO, approvals, chargeback statements, priority support, and up to 25 million requests per month.
VPC or self-hosting, DPA, custom routes, reconciliation ingestion, and SLA.
Talk to salesMachine-readable terms: pricing.md. Plans are subject to contract and current product availability.
Evidence library
A practical explanation of the proxy layer between applications and model providers, including routing, caching, policy, and observability.
Read the guide →How to route routine requests to lower-cost models without silently sacrificing output quality.
Read the guide →A plain-language guide to the FrugalGPT research and how to translate cascades into production controls.
Read the guide →What RouteLLM measures, what its results mean, and how to avoid copying benchmark claims into production forecasts.
Read the guide →How semantic response caches work, when they save money, and the controls needed to prevent unsafe reuse.
Read the guide →A side-by-side guide to provider prefix caches and gateway response caches.
Read the guide →Questions
An LLM gateway is a control layer between applications and AI providers. FrugalAI keeps the API OpenAI-compatible while centralizing model routing, caching, metering, budgets, and decision records.
FrugalAI routes eligible requests to the lowest-cost model that meets a defined quality bar, reuses approved repeat responses, and applies budget policy before spend occurs. Monitor mode measures the opportunity on your traffic before enforcement.
No. FrugalAI brokers requests to model providers. You keep provider choice while gaining one policy, cost, and audit layer.
Yes. Monitor mode records what the router and cache would have done while the existing provider path continues to answer. Teams move into enforcement only after reviewing their own evidence.
Start with a counterfactual, not a contract
Connect one workload in monitor mode. FrugalAI records the route it would choose, the quality checks required, and the estimated savings without changing the production answer.
Plan a monitor-mode pilot