Transparency and optimization for your
LLM traffic.
Your teams run AI in a dozen places. bttrfly enables you to see spend by team and app, enforce budgets and policy, and cut costs with a quality guarantee.
You can't govern what you can't see.
Your AI spend arrives without a breakdown
LLM calls leave from a dozen apps, keys, and providers, so the invoices land as totals with no team or application attached. Budgets and policy have nowhere to live, and the question of what a given team or app actually costs stays open.
And most of what you pay for is noise
Once you can see the traffic, you can see how much of it never needed to be sent. Stale history, duplicated blocks, and raw tool logs are billed on the way in and on the way out, and they crowd out what the model should be paying attention to.
See the traffic, and cutting it becomes the easy part.
Worried that compression makes quality worse? That is the right instinct, so we prove the opposite. Every workload runs against a control group, and we publish the change in quality next to the savings. On the workloads below, compressed context scored higher rather than lower.
Every call, attributed to a team and an app.
Before anything gets optimized, you get the picture you don't have today: where your AI spend comes from, which application produced it, and how it moves month to month.
See roughly what you'd save.
Savings depend on what you send, not just how much. Pick your heaviest workload and your monthly spend, and this shows the median we measured against a control group. We confirm your real number with a free audit on your own traffic.
Estimate only, from medians measured against a control group. Short, clean prompts have little headroom, and we don't touch what we can't improve.
Everything it takes to run LLM traffic across an organization.
Compression and caching are the easy part, and open source does them for free. What you buy above them is the quality guarantee, the governance, and the compliance and support that make it safe to put every team's traffic in one place. We build and run all of it.
Every number here was measured, not estimated.
Every optimization runs against a control group, so you see the real savings and any change in quality for each workload. When something doesn't pay off, we tell you and switch it off.
| Workload | Tokens | Quality Δ | Status |
|---|---|---|---|
| RAG and long-document Q&A | −61% | +3.1% | The biggest wins live here |
| Coding agents | −48% | +1.2% | Code-aware routing |
| Agentic, tool-heavy work | −54% | +2.0% | Cleaner tool output |
| Short, clean prompts | ±0% | ±0% | No real headroom, so we leave them alone |
See what your LLM spend actually delivers.
Token counts show consumption, and your dashboard should show results.
Cost per team tells you who spends the money, not what the money achieved. Where a workload has a measurable outcome, such as coding agents, agentic pipelines, and retrieval systems whose answers can be scored, bttrfly puts the spend and the result side by side. The dashboard answers what you got for the money, not just who spent it.
We show outcome metrics only where outcomes are measurable. That starts with engineering and agentic workloads, and extends one workload at a time. Everywhere else you still get the spend, attributed by team and application.
How we measure →| Team or app | Spend | What it delivered |
|---|---|---|
| Coding agents | €18,400 | 312 merged pull requests, about €59 each |
| Support answer assistant | €7,200 | 14,800 resolved tickets, about €0.49 each |
| Agentic data pipeline | €5,100 | 92% of runs finished without a human step |
| General internal chat | €4,300 | No reliable outcome signal yet, so we show spend only |
Transparency and optimization for your LLM traffic.
We start with an audit of your real traffic, not a contract.