About SOTAAZ Lab

SOTAAZ Lab tests local AI models and the tools around them on its own hardware, and writes down what was measured, how, and what it means for a choice you have to make. Posts are signed by the lab, not by a person. The point is that you can check the method yourself.

What we measure on

Two NVIDIA A100 80GB PCIe cards and an AMD EPYC 7742 64-core CPU. Some posts measure a hosted model through its API instead; those posts say so.

We have started adding a “Tested on” box under the title of measured posts, a few posts at a time: GPU, driver and CUDA version, runtime and build, and the date the numbers were taken. When a post mixes setups or quotes numbers measured elsewhere, the box also says which numbers it covers. We fill in only what the run logs, the setup notes in the published post, or files committed at measurement time (the dependency lock file, the figure script) confirm. An empty field means we could not confirm it.

We write the rule down before we score

Before scoring, we write down what we will compare and how large a difference has to be to count, and we commit that file. In the GPT-2 ladder series the plan is committed before the runs, and the download includes it with its commit times, so you can check the order. When part of a plan was written after some results were already in, the plan says which part.

Raw logs and reproduction packages

Three measured posts have a reproduction package you can download from the post itself, no account needed: the llama.cpp KV cache quantization benchmark, Strata vs llama.cpp on the same model file and the GPT-2 noise floor. Each zip holds the harness, the raw logs, a results CSV recomputed from those logs, and a README with the commands to rerun it.

Two code kits are also free with a free account:

The other measured posts don't have a package or kit yet. Their setup is described in the post itself.

AI helps us write; we run the measurements

We draft and edit posts with AI models. The measurements are not generated. We run them on the hardware above and check the numbers in the post against the logs before publishing. When a post quotes someone else's number, it says where the number came from.

When a number is wrong

We fix it in the post, put a dated note at the top saying what was wrong, and mark the corrected section. We don't swap numbers silently. When the software we measured changes, we re-run it and add a dated update next to the old result instead of deleting it. For example, the TurboQuant posts carry an August 2026 correction, and the Jeff vs Jev post carries a September 30, 2026 re-measurement on Jeff v1.1.

If you find a wrong number, write to contact@nestoz.co.