Cost, tokens and speed: the live-demo measurement
When and where. 13 September 2026, on our public demo at demo.trymithrandir.com: a real open-source codebase with a fictional company's conversations, tickets and decisions laid over it. Hundreds of entries: questions about the code and the decisions behind it.
The two sides. Each question was answered twice, in parallel. Once by an agent that asks Mithrandir. Once by a coding agent with the same model and read access to the same repository, working the way coding agents work today: searching and reading files until it can answer. Same model on both sides (gpt-5-mini), same questions.
What was counted. Total tokens on each side, input and output, across every step it took. Total time from the question being asked to the final answer arriving.
How it was summarised. For each question we take the ratio between the two sides, then report the median of those ratios, so one unusually long question cannot carry the result. That gives 14× fewer tokens and 2.2× faster (a median of 8.2 s against 18.6 s).
How it was priced. Tokens are turned into dollars at Claude Opus list prices ($5 per million input tokens, $25 per million output), using each side's real split between input and output. We use Opus list prices because that is the model heavy agent users commonly run; the model that actually ran costs less, so read the dollars as the ratio made concrete, not as what the run cost. That gives $0.031 against $0.234 per answer: 8.3× lower, an 88% smaller bill.
Read it with this in mind
The codebase is our demo, not yours. The coding agent reads the repository only; it cannot see the conversations and tickets Mithrandir has read, which is the situation the product replaces, and also why this run says nothing about answer quality. Quality is measured separately, below.
