Where every number on our homepage comes from.

Every number on our homepage is listed here: where it comes from, how it was measured, what the other side of the comparison was, and what the measurement cannot tell you. If you want to run the same comparison on your own repository, that is what the first call is for.

Measured, not estimated

An eighth of the cost. Every answer sourced.

8.3 times
lower cost$0.031 vs. $0.234 per output at Claude Opus list prices, against an agent mirroring the request with the exact same tools, minus Mithrandir. An 88% smaller AI bill, for better results.

Without Mithrandir$0.234

With Mithrandir$0.031

14 times
fewer tokens14× fewer tokens across the median of hundreds of entries. Coding agents re-read the repository turn after turn; Mithrandir retrieves once.
2.2 times
faster resultsMedian 8.2 s vs. 18.6 s from request to result, same questions, same setup on both sides. Faster, cheaper, better results with Mithrandir.
100 percent
of answers grounded in external sourcesNot a single output returned without exact sources out of hundreds of trials.
1.5 times
better answersRated half again as good as the same model working alone, on the same questions, graded without knowing which side was which.
zero
leaks of contextOut of hundreds of trials and thousands of customer entries, not a single leak of context.

Cost, tokens and speed: the same questions on our public demo, answered by an agent using Mithrandir and by a coding agent with the same model and the same repository, without it. Medians of per-question ratios; dollars at Claude Opus list prices. Quality and sourcing come from our pilot.

At a glance

Claim on the homepageWhere it comes from
8.3× lower cost · 88% smaller billLive-demo measurement
14× fewer tokensLive-demo measurement
2.2× faster resultsLive-demo measurement
100% of answers sourcedPublic demo and pilot
1.5× better answersPilot, blind-judged
0 leaks of contextPilots, demo and internal audit
Savings calculatorEstimate, not a measurement

Cost, tokens and speed: the live-demo measurement

When and where. 13 September 2026, on our public demo at demo.trymithrandir.com: a real open-source codebase with a fictional company's conversations, tickets and decisions laid over it. Hundreds of entries: questions about the code and the decisions behind it.

The two sides. Each question was answered twice, in parallel. Once by an agent that asks Mithrandir. Once by a coding agent with the same model and read access to the same repository, working the way coding agents work today: searching and reading files until it can answer. Same model on both sides (gpt-5-mini), same questions.

What was counted. Total tokens on each side, input and output, across every step it took. Total time from the question being asked to the final answer arriving.

How it was summarised. For each question we take the ratio between the two sides, then report the median of those ratios, so one unusually long question cannot carry the result. That gives 14× fewer tokens and 2.2× faster (a median of 8.2 s against 18.6 s).

How it was priced. Tokens are turned into dollars at Claude Opus list prices ($5 per million input tokens, $25 per million output), using each side's real split between input and output. We use Opus list prices because that is the model heavy agent users commonly run; the model that actually ran costs less, so read the dollars as the ratio made concrete, not as what the run cost. That gives $0.031 against $0.234 per answer: 8.3× lower, an 88% smaller bill.

Read it with this in mind

The codebase is our demo, not yours. The coding agent reads the repository only; it cannot see the conversations and tickets Mithrandir has read, which is the situation the product replaces, and also why this run says nothing about answer quality. Quality is measured separately, below.

Every answer sourced

Across hundreds of trials on the public demo and in the pilot, every answer cited at least one record from a connected source: a thread, a ticket, a document, a commit. A citation shows where an answer came from, not that it is right; that is what the next measurement is for.

Better answers: the blind-judged pilot

The two sides. In the pilot, every question asked of Mithrandir was also put to the same model on its own, with no Mithrandir behind it.

The judge. A separate model call compared the two answers without knowing which was which: labels removed, order shuffled. It scored each from 0 to 5 on three things: whether the answer is grounded in real sources, whether it answers the question directly, and whether it is complete. Mithrandir's answers scored half again as high: 1.5×.

Read it with this in mind

The judge is from the same model family as the answers, and the pilot is small. Read it as “the same model answers better with the memory behind it”, not as a ranking against other products.

Zero leaks of context

No answer has ever included something its asker was not cleared to see: not in any pilot, not on the public demo, and not in our internal access-control audit (August 2026). Before that, an adversarial drill ran 91 tests against permission boundaries; all 91 held.

Read it with this in mind

Zero observed is not proof of zero possible, and the audit was ours, not a third party's. That is why access is enforced before anything reaches the model, rather than by asking the model to be careful. The security overview describes what that means in practice.

The savings calculator is an estimate

The calculator on the homepage is a model, not a measurement. It borrows one measured number, the 88% saving, and estimates the rest. It starts from a heavy-usage AI bill for a 15-person team in each sector, built from usage figures published by Anthropic and industry reports, at Claude Opus list prices.

It then scales that bill to your team size slightly faster than headcount, because a bigger company has more material and an agent reads more of it before it answers. That growth rate is our derivation, not a measurement. Finally it applies the 88% saving from the live-demo measurement above, to every sector and every size alike.

Read it with this in mind

Treat the result as an order of magnitude, not a quote. Your own number depends on how your people use agents today, which is exactly what we measure with you in the first weeks.

Measure it on your own work.

Bring one workflow where context keeps getting lost. We will run the same comparison on it with you, and show you the raw numbers.

Book a call