In this post we introduce examroom, a new open-source Python library for testing how far you can push local language models. What is a local model? A local model is one you can run on your own hardware (for example a DGX Spark) instead of sending your data to a cloud service.
Why we built it
The library came out of one of our projects, where language models help doctors make decisions. We are working on colorectal cancer (CRC): we use historical data to train machine learning models that predict how likely a patient is to need surgery, given how their treatment is going. A language model then presents those predictions to the doctor in plain words.
One rule defines the whole project: data and models must stay local. No patient data can leave the hospital.
That makes open models such as Qwen or Gemma very interesting. They are free to download, they perform surprisingly well for their size, and they run on reasonably priced hardware. The question is: are they good enough for our use case?
Public benchmarks only answer part of that question. They tell you how good a model is at coding or maths, but not how it will handle a very specialised task like ours. There is a second problem too: to run on local hardware, models are often compressed (the technical term is quantisation), and we need to know how much ability is lost along the way.
So we built examroom: a library to write your own exam for language models, based on your own task.
examroom: an exam for your task, and a record of every answer
With examroom you write a small set of test questions about the one job you care about, and then give the same exam to as many models as you like. Every answer to every question is saved, so you can go back and read it later.
Most benchmarks ask "which model is best?". examroom asks something more practical: "which is the smallest model that is good enough for this job, and what exactly does it get wrong?" A smaller model means cheaper hardware and faster answers, so it is worth knowing where the line is.
An exam is a folder you can read
An exam (we call it a suite) is a folder of plain text files, one question per line. A question can be as simple as this:
"The car wash is fifty metres away, so I will…" → options: walk or drive.
There are three kinds of question:
- Multiple choice. We look at which option the model finds more likely. This works even on very small models that cannot follow instructions yet.
- Free-text answers. The model writes an answer, and a small program shipped with the exam checks it.
- Tool use. The model gets a set of tools (for example "find a patient" or "predict the risk") connected to a simulated service. It can use them until it is ready to answer, and we score everything it did along the way.
Large exams are not meant to be typed by hand: the idea is to generate them with the most capable models available.
When an exam is ready, you freeze it. From then on it cannot be edited: any change becomes a new version. This guarantees that two results on the same exam are always comparable.
Every answer is saved
For each run, examroom records which model was used and with which settings, how fast it was, how much memory it needed, and, for every single question, the score, the model's answer and the reason for the score.
Because everything is saved question by question, you can compare models in detail:
- see exactly which questions one model gets right and another gets wrong, with a statistical check that tells a real difference apart from luck;
- follow a single question across every model you have tested;
- browse all results in a single web page.
Any model, under the same rules
The same exam can be given to models running on your own machine, to models served by providers such as OpenRouter or OpenAI, and to Claude. Two things keep the comparison fair:
- Same settings for everyone. Models are only ranked together when they ran under the same conditions, for example with "thinking" switched off and the same limit on answer length.
- Size classes. examroom looks up the size of each model and places it in a class: tiny (up to 1 billion parameters), small (up to 15 billion) and local (anything that fits in 128 GB of memory). The class tells you where the model can actually be installed.
"Good enough" means passing every part
An average score can hide the one skill your application cannot do without. So examroom asks you to set a pass mark for each task, for example 95% for picking the right tool and 90% for everything else. It then lists the models by size and names the smallest one that passes every task.
The pass mark is strict. A model that gets 92 out of 100 has not proven that it is above 90%: with only 100 questions, a couple of those may be luck. examroom takes this margin of error into account, and a model passes only if it is above the mark even in the unlucky case.
Tool use, judged on the whole conversation
For tasks where the model has to use tools, the exam includes a simulated copy of the service the model will use in real life, filled with made-up data. examroom lets the model call the tools, gives it the results, and repeats until the model writes its final answer. The score depends on the whole exchange: which tools were used, with which inputs, what came back, and what the model finally told the user.
A leaderboard that cannot be gamed
If exam questions are public, future models may simply learn them by heart. So examroom splits each exam into a public half and a hidden half. The hidden half is kept private, and our public leaderboard ranks models on it. Only the scores are published, never the hidden questions or answers.
Each model also takes the public half, and the leaderboard shows examples of what it got wrong: the question, the expected answer, the model's answer and why it was marked down. In one example, a tumour 2.5 cm from the anal verge was classified as "middle" rather than "low".
The code that generates our questions is public, but it relies on a secret key, so nobody can use it to recreate the hidden questions.
For the technically curious, a typical session takes four commands:
examroom freeze suites/crc_clinician_v1.0
examroom run suites/crc_clinician_v1.0 openai:qwen/qwen3.5-9b \
--opt base_url=https://openrouter.ai/api/v1 --opt api_key_env=OPENROUTER_API_KEY \
--profile nothink@1
examroom select suites/crc_clinician_v1.0 --profile nothink@1 # smallest model that passes
examroom diff <run_a> <run_b> # what changed, question by question What we learned on our medical use case
We wrote an exam for our CRC project. It tests the language model that sits between the doctor and the prediction tools: the doctor writes a request in plain language, and the model has to understand it, use the right tool, and explain the result faithfully. The exam is in English and Italian, because the system will be used in Italian hospitals. All patients and reports in it are made up: no real patient data is involved. The questions are meant for preliminary testing only and have not yet been reviewed by clinicians.
The headline result: the smallest model that passes every task is Gemma 4 26B-A4B. No model in the "small" class (15 billion parameters or fewer) passes, even though one of them has an average score of 96.6%.
Which model size is enough
We tested 14 open models on the hidden half of the exam: 1,000 questions in English and Italian. The exam covers the five things the language model does in the project:
- Extract: turn an MRI or lab report into a structured patient record.
- Route: pick the right tool for a doctor's request, or decline the request.
- Next turn: decide what to do next in the middle of a conversation.
- Narrate: explain the results of a tool without adding numbers of its own.
- Grounded answers: answer a question from the hospital protocol, citing the relevant passages.
Five tasks in two languages make ten tests. All models ran with "thinking" switched off, through OpenRouter. The pass marks were 95% for routing and 90% for everything else.
| Model | Size (parameters) | Class | Average score | Tests failed (of 10) |
|---|---|---|---|---|
| Llama 3.2 1B | 1.24B | small | 20.1 | 10 |
| Llama 3.2 3B | 3.21B | small | 61.7 | 10 |
| Gemma 3 4B | 4.3B | small | 77.9 | 9 |
| Ministral 8B | 8.92B | small | 89.0 | 10 |
| Qwen3.5 9B | 9.65B | small | 96.6 | 5 |
| Gemma 3 12B | 12.2B | small | 92.7 | 8 |
| Ministral 14B | 13.9B | small | 94.0 | 6 |
| Mistral Small 3.2 | 24B | local | 95.9 | 3 |
| Gemma 4 26B-A4B | 25.8B (about 4B active) | local | 98.5 | 0 |
| Gemma 3 27B | 27.4B | local | 94.2 | 6 |
| Qwen3.8 27B | 27.8B | local | 99.2 | 0 |
| Qwen3 30B-A3B | 30.5B (3.3B active) | local | 89.7 | 8 |
| Gemma 4 31B | 31.3B | local | 99.6 | 0 |
| Qwen3.8 Flash | 180B | above local | 99.4 | 1 |
Four things stand out.
The average is the wrong number to look at. Qwen3.5 9B has a higher average than Mistral Small (24B) and Gemma 3 27B, yet it fails five tests. An application breaks at its weakest step, not at its average.
Italian is where the smaller models fall. Four of Qwen3.5 9B's five failures are in Italian, and so are two of Mistral Small's three. An English-only exam would have approved models that would not work in an Italian hospital.
Size tells you little. Gemma 4 26B-A4B passes everything while using only about 4 billion of its parameters for each word it writes. Qwen3 30B-A3B, which is built in a similar way, fails eight tests. Gemma 3 27B, the previous generation at the same size, fails six. How recent a model is matters more than how big it is.
The bar is strict on purpose. Qwen3.8 Flash gets 99 out of 100 on Italian routing and still fails. With only 100 questions, 99% is not enough to be confident that the true score is above 95%. The fix is more questions per task, not a lower bar.
For our project, a model in the 26–31B range is enough on this exam, and nothing at 15B or below is, for now. The top three models are within about one point of each other, so telling them apart will take a harder exam.
Where two models differ, question by question
Gemma 3 12B scores 3.9 points lower on average than Qwen3.5 9B. Because examroom saves every answer, we can see where the gap comes from:
| Test | Qwen3.5 9B | Gemma 3 12B | Difference | Real difference? |
|---|---|---|---|---|
| Extract (English) | 97.5 | 91.4 | -6.1 | not clear |
| Extract (Italian) | 95.4 | 93.2 | -2.1 | not clear |
| Grounded answers (English) | 95.0 | 89.0 | -6.0 | not clear |
| Grounded answers (Italian) | 92.0 | 82.0 | -10.0 | not clear |
| Narrate (English) | 96.7 | 100.0 | +3.3 | not clear |
| Narrate (Italian) | 92.0 | 97.3 | +5.3 | not clear |
| Next turn (English) | 100.0 | 93.0 | -7.0 | yes |
| Next turn (Italian) | 99.7 | 93.0 | -6.7 | yes |
| Route (English) | 100.0 | 91.5 | -8.5 | yes |
| Route (Italian) | 98.0 | 96.5 | -1.5 | not clear |
| Total | 96.6 | 92.7 | -3.9 |
The gap is in choosing what to do. Gemma 3 12B gets 8 English routing questions and 12 next-turn questions wrong that Qwen3.5 9B gets right, and those differences are too large to be luck. Its explanations, if anything, are slightly better. Every one of these questions can be opened to read the two answers side by side.
Using tools, from start to finish
Our second exam is an early draft with 240 questions. It gives the model a simulated copy of the surgery-risk service our project will offer. The service has six tools:
- find a patient;
- read the patient's clinical history;
- predict the probability of surgery, including "what if" scenarios;
- explain which factors drove a prediction;
- find the most similar past patients;
- read the documentation of the prediction model.
The model has to collect what a medical team needs through these tools, and report it without inventing numbers. There are three tasks, each in English and Italian:
- Analysis: the prediction, its main drivers and similar patients for one patient;
- What-if: the patient's prediction, plus the prediction under a hypothetical change;
- Limits: requests the tools cannot fully answer, such as missing data, an unknown patient, or an outcome the service does not predict.
| Model | Total | Analysis | What-if | Limits |
|---|---|---|---|---|
| Qwen3.8 27B | 97.7 | 100.0 | 97.5 | 95.7 |
| Gemma 4 31B | 96.6 | 99.8 | 97.5 | 92.5 |
| Qwen3 30B-A3B | 93.8 | 96.0 | 99.3 | 86.2 |
| Ministral 8B | 91.1 | 89.8 | 99.6 | 83.8 |
| Ministral 14B | 89.2 | 89.3 | 99.6 | 78.8 |
| Qwen3.5 9B | 88.4 | 95.0 | 87.1 | 83.1 |
| Gemma 4 26B-A4B | 87.8 | 88.6 | 94.9 | 80.0 |
| Mistral Small 3.2 | 82.4 | 80.7 | 95.0 | 71.4 |
| Gemma 3 12B | 2.1 | 0.0 | 1.1 | 5.2 |
Admitting what the tools cannot tell you is the hard part. Every model scores lowest on "limits". Typical mistakes:
- presenting the probability of surgery as if it were the risk of a permanent stoma;
- making up a patient ID instead of searching for the patient;
- calculating an average follow-up time that no tool ever returned.
The ranking changes. Ministral 8B is fourth here, ahead of Gemma 4 26B-A4B, which passed everything on the first exam. Following written instructions in a single reply and using real tools over several steps are different skills.
Test the setup you will actually use. Gemma 3 12B scored 92.7 on the first exam and 2.1 here. In 215 of 240 conversations it described the tool calls in its reply instead of actually making them. Three of the smallest models (Llama 3.2 1B and 3B, Gemma 3 4B) could not be tested at all, because no provider on OpenRouter offers them with tool support.
One example shows why we judge the whole conversation and not just the final reply. The doctor asks: "For patient P-9268: what is the probability of surgery within two years, and how would it change if the CEA after treatment were 320 ng/mL?" (CEA is a blood marker used to monitor the tumour.) The service only accepts CEA values up to 100.
Gemma 4 31B asks the service for the current prediction (57%), then asks for the scenario with CEA at 320. The service replies with an error: the value must be between 0 and 100. The model tells the doctor: "I cannot calculate the change for a CEA of 320 ng/mL because the model only accepts values between 0 and 100."
Qwen3.5 9B asks for the current prediction (57%), then asks for the scenario with CEA at 32 instead of 320. The service replies 89%. The model tells the doctor: "If the CEA were hypothetically 320 ng/mL, the probability would increase to 89%."
Seven of the eight models that used the tools behaved like Gemma 4 31B. Qwen3.5 9B quietly sent 32 instead of 320, and then presented the result as the answer for 320. Its reply reads well, and the 89% really did come from the service. The mistake is visible only in what the model asked behind the scenes. We score this conversation zero, because the number describes a scenario the doctor never asked about.
Conclusion
Can a local model do the job? For our project the answer is yes, but not just any model. On our exams a recent model in the 26–31B range, small enough to run inside a hospital, handled the tasks we need, while the smaller ones did not yet. On tool use, the two best models were Qwen3.8 27B and Gemma 4 31B.
Beyond our own use case, a few lessons apply to anyone choosing a language model:
- Write an exam for your own task. Public rankings did not predict which models would work for us.
- Look at the weakest step, not the average. A model with a great average can still fail the one step you depend on.
- Test in the language of your users. Several models that looked fine in English failed in Italian.
- Test the real setup. A model that answers questions well may still be unable to use tools.
- Read the answers, not only the scores. The most dangerous mistakes are the ones that sound right.
examroom is open source. You can find the code on GitHub and the results on the public leaderboard. If you are trying to work out whether a local model is good enough for your own task, we would love to hear how it goes.