All Posts

Give the Intern a Budget

by Rick Radewagen2 min read

A fixed level of machine intelligence gets about ten times cheaper every year. The frontier gets more expensive, because it does more work per answer. Almost every team watches one curve and ignores the other, and that is why AI analytics bills surprise people.

The question is no longer what AI analytics costs. It is which questions deserve which level of intelligence. Run everything on the frontier and you get a bill nobody can explain. Run everything cheap and you quietly lose the answers that needed more thought.

Most questions do not need the frontier

Most real questions are simple lookups against a well-modelled domain. A cheaper model answers them correctly, faster, for a fraction of the cost, and users notice the speed more than anyone notices the savings. A minority genuinely need the frontier: multi-step investigations, ambiguous questions, decisions where being wrong is expensive. Set an economical default and escalate deliberately.

There is a test hiding in this. If a domain only works on the most expensive model, that is usually a fact about your data model, not about the questions. Clean context makes cheap intelligence sufficient. When the economical mode handles a domain, the context is in good shape.

Budgets make people choose

Give people budgets. It sounds unkind and it is not. An analyst making pricing decisions and an intern exploring should not have the same allowance for delegating work to a machine, for the same reason they do not have the same expense limits. And a ceiling changes behavior. With no limit, everything looks equally worth asking. With one, people spend on what matters to them, which is exactly the allocation you wanted and could never have made centrally.

Expect a power law. In every deployment we see, a handful of people drive most of the usage. They are usually your highest-value users, not your problem. Talk to them instead of emailing everyone about responsible usage.

Rising spend is usually adoption

If the bill went up because forty more people started getting answers instead of waiting three days for them, the purchase is working. The number to watch is not total spend. It is whether spend lands on work worth doing.

Real waste looks different: a scheduled report nobody reads, or one badly scoped conversation looping for twenty minutes because the agent is missing context. Neither is fixed by a cap. The first needs an audit. The second needs better context, which is the same work that makes everything else cheaper.

When spend climbs, ask what it bought before you ask how to cut it. Cost here is a governance problem, not a bill you receive.

If this excites you, we'd love to hear from you. Get in touch.

Rick Radewagen

Rick is a co-founder of Dot, on a mission to make data accessible to everyone. When he's not building AI-powered analytics, you'll find him obsessing over well-arranged pixels and surprising himself by learning new languages.