Pay as you gain · zero risk
TokenomiCo sits between your app and the API you already use and compresses what repeats on every call — history, tools and instructions. In a real 20-turn agentic session we measured 28% lower input cost (and on a single short prompt we are honest: there are no savings). You only pay 40% of what you save: no savings, no invoice. No subscription, no contract, not a single line of code to change.
the model
We don't sell subscriptions, we sell results. We measure every token you didn't spend and split the gain — the bigger share stays with you. If a month brings no savings, your invoice is zero. You can't lose.
how it works
No migration, no rewrites, no model switch. If you have an AI app, you can connect today.
Point the SDK you already use at the TokenomiCo endpoint. Same API, same models, same code. Your provider key stays encrypted and under your control.
Our technology trims your prompts in real time while preserving meaning. And when a request has no real gain, it passes through untouched — you're never worse off.
Every response shows how much you saved. At the end of the month, a statement details the savings — your invoice is 40% of it, and the other 60% stays in your pocket.
benchmarks · measured, not estimated
Every compression number below came from the compiler actually running (LuaJIT 2.1), deterministic and reproducible. And we're explicit about what is a direct measurement, what is a calibrated model and what we haven't measured yet.
Compression by domain (direct measurement)
| domain | tokens | reduction |
|---|---|---|
| Ops / incident response | 106 → 56 | −47.2% |
| Code refactoring | 98 → 54 | −44.9% |
| Legal audit (LGPD) | 117 → 66 | −43.6% |
| Data analysis | 98 → 56 | −42.9% |
| CI/CD pipeline | 104 → 62 | −40.4% |
| Security recon | 119 → 71 | −40.3% |
| TOTAL | 642 → 365 | −43.1% |
A narrow range (40–47%) means robustness: no "lucky" domain carrying the average.
Agentic loop: what if γ is worse?
The history-compression factor (γ) is the parameter with the most uncertainty — so we publish the whole range, not just the best case:
| γ (history) | savings @10 steps | @30 steps |
|---|---|---|
| 0.10 · aggressive | 94.8% | 93.8% |
| 0.20 | 91.2% | 88.0% |
| 0.30 · base | 87.7% | 82.1% |
| 0.40 · conservative | 84.1% | 76.3% |
Even in the worst case (γ=0.40), loop savings stay above 76%. @ $3/MTok, 30 steps: $0.98 → $0.18.
run the numbers
Drag the sliders with your own numbers and the receipt shows, instantly, how much would come back to you every month.
In our tests, natural-language prompts shrink by up to 30–40%. Be conservative if you like — even in the worst case (0%) you pay nothing.
savings statement · simulation
— you're only billed on real gains —
for developers
A true drop-in: one URL and one header. You control compression per request and see the savings on every response.
curl https://tokenomico.com/v1/anthropic/proxy \ -H "authorization: Bearer tk_live_..." \ -H "x-toke-compress: on" # on|off per request \ -H "content-type: application/json" \ -d '{"model":"claude-sonnet-5","max_tokens":300, "messages":[{"role":"user","content":"..."}]}' # response includes: x-toke-compressed: true x-toke-tokens-saved-est: 1284
stream:true responses pass through byte by byte, with real usage captured from the SSE for billing.
A proxy for provider API endpoints. No subscriptions, no third-party chat apps.
Monthly proof with requests, original vs. sent tokens and the savings in dollars.
no fine print
On tasks highly sensitive to exact wording, it can. That's why it's
optional per request (x-toke-compress header): turn it on for tolerant
workloads (classification, RAG, long context) and off for extraction/JSON and critical ones.
Short prompts and structured content are passed through uncompressed automatically. And when there's no
real gain, the original text is sent — you're never worse off than before.
Zero invoice. The model is 40% of measured savings, not of your usage.
The proxy processes text in transit to compress it (TLS in and out) and stores no prompts or responses — only token metrics. One nuance, for transparency: the compressed rewrite of recurring passages (never the original) may be retained for up to 7 days to speed up subsequent turns. Your BYOK key is encrypted at rest (AES-256-GCM).
Monthly: the invoice closes on the 15th and arrives by e-mail with the savings statement (PoC) and a Stripe payment link. Grace period until the 18th; after that the proxy is suspended (HTTP 402) until settled — pay and you're back instantly. New accounts get a grace cycle: your first invoice only comes the month after you sign up. Your data and keys remain yours.
Anthropic, OpenAI, Gemini, Grok (xAI), Perplexity, DeepSeek, Mistral, Groq, Kimi (Moonshot), Together AI and OpenRouter — one TokenomiCo key for all of them.
start now — no card required
Create your free account, connect your key and see your first savings statement today. The only scenario where you pay is the one where you've already gained more.