Solutions · model routing

The right model for every call. Automatically.

Router sits in the Caveman Platform request path and sends each call to the cheapest model in your pool that passes your evals. If nothing cheaper passes, nothing moves.

What Router does
the rule
the cheapest model in your pool that passes your evals
the fallback
nothing cheaper passes — traffic stays on the pin you already run
the rollout
record → replay → shadow → canary → active
on regression
the route rolls back on its own
the claim
inferred until provider-causal · verified savings start at $0.00

Router and Caveman Platform are in private development. Savings stay inferred until they are provider-causal.

01Problem

Every call goes to the priciest model. Switching is a risk nobody owns.

01
One model, pinned

Teams pin the model they have watched work, and every call in the application goes to it.

02
Proving a swap is work

Building the harness, running the comparison, owning the regression — nobody wants that on their name.

03
So easy calls pay top rate

The hard requests and the trivial ones are billed at exactly the same price.

02Mechanism

The cheapest model that passes. Or traffic stays put.

the rule
Router sends each call to the cheapest model in your pool that passes your evals. If nothing cheaper passes, traffic stays put.
eval-gated rollout
Record → replay → shadow → canary → active. Each stage has a gate, and a route only reaches live traffic after it clears the stage before it.
automatic rollback
A regression rolls the route back without waiting for anyone. The route you were serving keeps serving.
what we will claim
Savings stay inferred until they are provider-causal. No invented percentage, and no dollar figure your own traffic has not produced.
  1. 01
    record
    real traffic, captured
  2. 02
    replay
    the candidate, on what already happened
  3. 03
    shadow
    alongside production, serving nothing
  4. 04
    canary
    a slice of live traffic
  5. 05
    active
    the route, in full
03Route map

One gate decides. The pool is yours.

your traffic, call by call

call
call
call
call
call
your evals
the switch, on every call
cheapest
fails your evals
next cheapest
routed here
your pin today
passes
most expensive
passes

your model pool, cheapest first

the cheapest model that passed takes the callnothing cheaper passes — traffic stays on the pin you already run

The map is the rule drawn, not a measurement: one lane is lit because it is the cheapest model that passed. Router and the Platform are in private development.

04Console

What actually moved. Including nothing.

Routing reports what it did, not what it hoped to do. In this demo workspace nothing has been rerouted yet — so the panel reads 0 rerouted, and the route pairs table lists what each provider served at catalog list price.

A route only shows up here after an eval-gated change moves a model.

Caveman Cloud routing view: 38.6M tokens across 8 models, 17.9% cache and compression, routing at 0.0% with 0 rerouted, $117.89 of catalog-priced route spend, an empty rerouted-traffic panel reading no downshift rows yet, and a route pairs table listing each provider and model with its request count and catalog subtotal.
routing · real product UI · demo workspace · private beta
05Products

What powers it.

Stop paying the top rate
for easy calls.

The managed plane

Router and Caveman Platform are in private development. Spend is priced from provider-reported usage against the public model catalog; unknown models stay unpriced, and verified savings start at $0.00 and move only on provider-causal evidence.