All Posts

Cafe Bench: how well can Dot run a business for a year?

by Anand Ani2 min read

The ultimate goal for Dot is to help our customers make better decisions, and hence create more value.

Most of the evals we've built so far measure pieces of that. Is the SQL correct? Does the chart answer the question? Did Dot hallucinate? While these do matter, they do not measure how good decisions made by Dot are.

Cafe Bench is our attempt to measure that more directly. We give Dot a small business with two years of history and ask it to run the business for a year. Every decision it makes has consequences, and those consequences show up in the data it reads next.

This week Anthropic released Claude Opus 5.5 and OpenAI released GPT-6 Sol and GPT-6 Luna. In celebration of this, we are making some results from the benchmark public.

Each model ran the same two simulated years (A and B). Left alone, the business goes from about $315k to $1.9M in net worth over the 455 days. Here is how that changes with each model in charge:

The data reveals something interesting: shelling out for a frontier model seems worth it. Would you not spend $50 extra to make $75k? It shows that the frontier can get a lot more expensive.

At its core, the business is a small coffee chain: four cafés in two districts, a rival café in each district, and 12,000 simulated regular customers, each with habits of their own. Dot starts with the chain having been running for two years.

Dot gets what a customer would give it. It has read-only SQL access to the company's own records: sales, queues and walkouts, staffing, stock, supplier terms, notices the business has received, and checks on the rivals' prices. It also gets a data dictionary and a short note from the owner. It does not see the customers' hidden preferences or anything about what the year holds.

It can change prices and quality, staffing for each part of the day, how much stock to keep, one-off ingredient orders, equipment (bought or leased), and whether to sign a take-or-pay supply contract. After each check-in it decides how long to wait before the next one, anywhere from 3 to 91 days. It gets 20 check-ins where it can make changes, plus up to six free wake-ups when a notice arrives.

The simulation then runs day by day. Decisions land in the same tables Dot queries, two days late, just like real reporting. Dot makes decisions for 365 days, then its last settings keep running for 90 more, so a decision whose cost arrives late still counts. The score is the chain's profit over all 455 days.

ModelYear AYear BAverageCostTime
Claude Opus 5.5 (medium)+$236,638+$216,454+$226,546$8.5722 min
GPT-6 Astra (medium)+$129,942+$184,064+$157,003$56.34132 min
Claude Fable 5.1 (medium)+$75,072+$204,709+$139,890$12.8127 min
GPT-6 Sol (medium)+$106,889+$90,213+$98,551$7.4128 min
GPT-6 Astra (low)+$87,107+$63,128+$75,117$22.6851 min
GPT-5.6 Sol (medium)−$19,374−$73,116−$46,245$9.4028 min
GPT-6 Luna (medium)−$22,696−$98,103−$60,400$1.8640 min
GPT-5.6 Luna (medium)−$226,377−$339,861−$283,119$1.1110 min

Things to keep in mind

  • Two runs per model. Ideally, the sample size would have been larger, but this is a first version and it's expensive to run.
  • The world is more nuanced. While we have done our best to make the simulations complex, they are obviously no match for the real world.

Cafe Bench was inspired by Andon Labs' Vending-Bench 2.

Anand Ani

Anand is a founding AI engineer at Dot. Builder at heart, part-time nerd, and a chess player on the side.