July 27, 2026 · Tokenless · 2 min read
Introducing Tokenless
The intelligent model router that sends each request to the rightsized model
Our router fans out your requests to multiple models in parallel and evaluates their thinking live. Your requests get routed to the cheapest model that is clearly on track to succeed at the job.
Agents are becoming a larger and larger proportion of AI spend, and yet most of what is spent on them is wasted. Between the turns that require deep reasoning and frontier intelligence, agents do a lot of mundane work such as reading files, summarizing short documents, or running tests. Not every one of these tasks needs to be sent to the most expensive models available.
Misunderstanding this idea has been disastrous for businesses in 2026. There were reports of Uber blowing through their entire yearly AI budget in just months, while Salesforce is freezing hires as AI spend explodes. An anonymous company accidentally spent $500M on AI. Tokenmaxxing has been a wildfire for productivity, but the burn is getting out of control.
This is why we’re launching Tokenless. Tokenless sends your prompts and agent turns to the right model each time, which maintains frontier performance while saving costs. It’s a drop-in replacement for OpenAI and Anthropic APIs, and you can bring your own keys to route traffic to your accounts. Give it a try in less than 5 minutes!
Actually choosing when to use which model is hard. It’s often guesswork based on so called “model astrology” — intuition and hearsay of how to use LLMs effectively, rather than hard evidence. Some say use Claude to plan and GPT to implement. Others say use GPT to plan and Claude to implement.
We made Tokenless to give a research-backed solution to this problem. We’ve collected hundreds of billions of tokens in model rollouts to understand how different LLMs react to different workloads. That data was used to train our custom foundation models to aid in routing.
We’ve also reimagined the algorithm behind solving the routing problem: we fan out requests to several models at once in parallel, watch them think, and keep the cheapest one that works, while cutting off the ones that don’t. We do this every turn so we’re always looking for chances to save.
We believe that while frontier models are pushing the capability of agentic work, lightweight open-weights models are pushing the viability of agentic work instead. Amazing model releases by Nvidia, Kimi, Deepseek, Minimax, and Thinking Machines have shown that open-weights models are capable of powering coding agents, chatbots, and copilots. By switching between frontier and open models, agents can get the best of both, and by releasing Tokenless, we’re giving everyone the tools to do so.
Stay tuned for a deep dive into the inner workings of our router and benchmarks coming out later this week!
Try routing some work through Tokenless and give us your honest feedback. Let us know what goes well, and what needs improvement.