Routers
The manager
Decides which brain and which worker gets each job, and how much to spend on it.
Every request has a cost in time and money. AI companies charge by the “token”, which is a piece of a word, and that includes everything sent along with the question, not only the answer.
So the manager’s job is to send each task to the cheapest brain that can do it properly, and to stop workers carrying far more than they need.
291,000 → 23,860
tokens for the same one-word request, before and after a fix
The managers I built
The task router
Works out what kind of job it is, picks a brain and a worker, runs it, checks the result against what was asked, and tries again or passes it up to a stronger brain if it failed.
The size router
Sends a simple message such as “thanks” to the small 8-billion model and everything else to a bigger one. Simple replies arrive in under a second instead of two.
Tool profiles
Six ready-made sets of tools: coding, code hosting, long runs, research, security practice, administration. A session starts with only the set its job needs.
The cost check
At the start of every real task, the worker asks whether a cheaper model could do it, and says so in one line.
The token diet
One tool compresses what gets sent to the outside models. Another chooses which notes to include, so a worker reads the three that matter, not all of them.
Switching in plain words
I can write “use the big one for three hours” in a chat. The model changes, and changes back by itself when the time is up.
What it taught me
- 1
Measure before believing. One of my tools was sending 291,000 tokens to ask a one-word question, because dozens of unused tools were attached. After the fix it sent 23,860.
- 2
Fewer tools in view is cheaper and safer at the same time.
- 3
A small fast brain for small things makes the whole system feel quicker than one big brain for everything.