Pay as you gain · zero risk

The same AI.
The same models.
At a fraction of the cost.

TokenomiCo sits between your app and the API you already use and compresses what repeats on every call — history, tools and instructions. In a real 20-turn agentic session we measured 28% lower input cost (and on a single short prompt we are honest: there are no savings). You only pay 40% of what you save: no savings, no invoice. No subscription, no contract, not a single line of code to change.

live · a real prompt shrinking
tokens 35 same answer, lower cost
AnthropicOpenAIGeminiGrokPerplexity DeepSeekMistralGroqKimiTogetherOpenRouter

the model

Pay as you gain: you only pay when you win.

We don't sell subscriptions, we sell results. We measure every token you didn't spend and split the gain — the bigger share stays with you. If a month brings no savings, your invoice is zero. You can't lose.

$0in subscriptions, quotas or minimums. Creating an account is free — forever.
$0invoiced in months with no savings. Our incentive is literally yours.
0 lock-inyour key stays yours (BYOK). Cancelling is reverting a base URL — no penalty.

how it works

Saving in 3 steps

No migration, no rewrites, no model switch. If you have an AI app, you can connect today.

1 — connect in minutes

Swap one URL, nothing else

Point the SDK you already use at the TokenomiCo endpoint. Same API, same models, same code. Your provider key stays encrypted and under your control.

2 — save on autopilot

Every request gets cheaper

Our technology trims your prompts in real time while preserving meaning. And when a request has no real gain, it passes through untouched — you're never worse off.

3 — verify and pay only the deal

Full transparency

Every response shows how much you saved. At the end of the month, a statement details the savings — your invoice is 40% of it, and the other 60% stays in your pocket.

benchmarks · measured, not estimated

Numbers from running code, not from slides.

Every compression number below came from the compiler actually running (LuaJIT 2.1), deterministic and reproducible. And we're explicit about what is a direct measurement, what is a calibrated model and what we haven't measured yet.

−43.1%tokens per turn, average across 6 real domains
direct measurement
−30.5% to −35.7%revalidated with a conservative BPE tokenizer — the floor, not the ceiling
cross-check
−28.3%input cost in a REAL 20-turn agentic session (Gemini in the cloud, compression + cache)
direct measurement
1.66 msper compile (~600/s single-thread) — invisible next to network latency
direct measurement

Compression by domain (direct measurement)

domaintokensreduction
Ops / incident response106 → 56
−47.2%
Code refactoring98 → 54
−44.9%
Legal audit (LGPD)117 → 66
−43.6%
Data analysis98 → 56
−42.9%
CI/CD pipeline104 → 62
−40.4%
Security recon119 → 71
−40.3%
TOTAL642 → 365
−43.1%

A narrow range (40–47%) means robustness: no "lucky" domain carrying the average.

Agentic loop: what if γ is worse?

The history-compression factor (γ) is the parameter with the most uncertainty — so we publish the whole range, not just the best case:

γ (history)savings @10 steps@30 steps
0.10 · aggressive 94.8%93.8%
0.20 91.2%88.0%
0.30 · base 87.7%82.1%
0.40 · conservative 84.1%76.3%

Even in the worst case (γ=0.40), loop savings stay above 76%. @ $3/MTok, 30 steps: $0.98 → $0.18.

Honesty about the gap: the 82% model assumes history compression at γ=0.30, which the current engine does not yet reach on dense prose — in the cloud, against real Gemini, we measured −28.3% input cost over 20 turns (compression + prefix cache; methodology and full series in the benchmark doc). On a single short prompt there are no savings — and the invoice is zero. Not yet measured: the effect of compression on answer quality per domain.

run the numbers

How much are you leaving on the table?

Drag the sliders with your own numbers and the receipt shows, instantly, how much would come back to you every month.

$2,000
30%

In our tests, natural-language prompts shrink by up to 30–40%. Be conservative if you like — even in the worst case (0%) you pay nothing.

for developers

Integrates in an afternoon? Try one coffee.

A true drop-in: one URL and one header. You control compression per request and see the savings on every response.

curl https://tokenomico.com/v1/anthropic/proxy \
  -H "authorization: Bearer tk_live_..." \
  -H "x-toke-compress: on"         # on|off per request \
  -H "content-type: application/json" \
  -d '{"model":"claude-sonnet-5","max_tokens":300,
       "messages":[{"role":"user","content":"..."}]}'

# response includes:
x-toke-compressed: true
x-toke-tokens-saved-est: 1284

Metered streaming

stream:true responses pass through byte by byte, with real usage captured from the SSE for billing.

APIs/REST only

A proxy for provider API endpoints. No subscriptions, no third-party chat apps.

Auditable statement

Monthly proof with requests, original vs. sent tokens and the savings in dollars.

no fine print

Straight questions, straight answers

Can compression change the model's output?

On tasks highly sensitive to exact wording, it can. That's why it's optional per request (x-toke-compress header): turn it on for tolerant workloads (classification, RAG, long context) and off for extraction/JSON and critical ones. Short prompts and structured content are passed through uncompressed automatically. And when there's no real gain, the original text is sent — you're never worse off than before.

What if I save nothing in a month?

Zero invoice. The model is 40% of measured savings, not of your usage.

Do you see my prompts?

The proxy processes text in transit to compress it (TLS in and out) and stores no prompts or responses — only token metrics. One nuance, for transparency: the compressed rewrite of recurring passages (never the original) may be retained for up to 7 days to speed up subsequent turns. Your BYOK key is encrypted at rest (AES-256-GCM).

How does billing work?

Monthly: the invoice closes on the 15th and arrives by e-mail with the savings statement (PoC) and a Stripe payment link. Grace period until the 18th; after that the proxy is suspended (HTTP 402) until settled — pay and you're back instantly. New accounts get a grace cycle: your first invoice only comes the month after you sign up. Your data and keys remain yours.

Which providers?

Anthropic, OpenAI, Gemini, Grok (xAI), Perplexity, DeepSeek, Mistral, Groq, Kimi (Moonshot), Together AI and OpenRouter — one TokenomiCo key for all of them.

start now — no card required

If you don't save, you don't pay.
That simple.

Create your free account, connect your key and see your first savings statement today. The only scenario where you pay is the one where you've already gained more.