We have published our first research report. It describes Treepai g0, a 110M-parameter language model we trained from scratch on one consumer GPU in 4.67 hours.
What we found
On five core benchmarks g0 averages 37.3%, against 38.6% for GPT-2 small. With our improved training recipe it reaches 38.9%, level with GPT-2 small. Against an early Pythia-160M checkpoint at a similar 2.1 billion training tokens, it leads on every core benchmark (37.3% against 28.1% on average).
In the same engine (llama.cpp, 8-bit) it was the fastest of the three models we tested, measured on consumer Apple-silicon hardware, with no measurable loss from 8-bit quantization.
It is not the best model of its size. SmolLM2-135M, trained on far more data, is well ahead on quality and is our target. g0 is also weak on story-level reading, where it trails GPT-2 small by 13 points.
Why it matters to us
We also tested a broader mix of training data. At this size it raised everyday common sense and lowered science, rather than improving both. That is consistent with our view that small models do better when they specialize. It is one test at one size, and we will repeat it on larger models.
What’s next
g0 is a research baseline, and its weights are not released. Our next, larger model, g1, is in training.
Read the report. Questions: research@treepai.com.