Research

Treepai g0

Research report · 7 October 2026

Report 01

Treepai g0: A 110M-Parameter Language Model Trained From Scratch on One GPU

Jose Franco · Treepai, Inc. · October 2026

Treepai g0 is our first model and the baseline for our later work. The report describes how it was built, how it was measured, and where it falls short.

Read the report (PDF, 24 pages)

You may download this report and share it unchanged, with credit to Treepai, Inc. All other rights are reserved.

Main results

  • Trained from scratch on 2.47 billion tokens in 4.67 hours on one consumer GPU. It averages 37.3% on five core benchmarks, against 38.6% for GPT-2 small and 39.0% for Pythia-160M. With our improved training recipe it reaches 38.9%, level with both.
  • Against an early Pythia-160M checkpoint at a similar 2.1 billion training tokens, it leads on every core benchmark (37.3% against 28.1% on average).
  • In the same engine (llama.cpp, 8-bit), it was the fastest of the three models tested: 4–32% more characters per second than GPT-2 small and 1.4–2.5 times SmolLM2-135M, with no measurable loss from 8-bit quantization. Speed was measured on consumer Apple-silicon hardware.
  • In a controlled test, replacing 40% of the educational training text with broad web text raised everyday common sense and lowered science, rather than improving both.

Where it falls short

  • SmolLM2-135M, trained on about 2 trillion tokens, is well ahead on every benchmark (49.7% on the core average). It is the strongest open model of this size and our long-term target.
  • g0 is weakest on story-level reading: 19.4% on LAMBADA against 32.6% for GPT-2 small. Its training data has almost no fiction.
  • Each full-length run was trained once. Differences under about 2 points are not meaningful at this size.
  • The data-mix test ran once per mix, and each mix also removed 40% of the educational text. Its conclusion is suggestive, not established. We will repeat it on larger models.
  • g0 is a base model with no instruction tuning or safety training. It is not suitable for any user-facing use, and its weights are not released.

How we measured

  • One evaluation harness for every model. The reference models were scored again with it, not copied from papers, after it reproduced GPT-2 small’s commonly published numbers.
  • The five core benchmarks were fixed before training. Four more were added afterwards and are reported separately.
  • Before training, we removed every training document that shared a 13-word sequence with a core benchmark question and its answer.
  • Final text-prediction scores use a sealed test set that was held out before training.
  • For changes to the training recipe, we measured run-to-run noise first and decided in advance what counts as a win: more than twice that noise.

The full method is in Sections 5, 7 and 9 of the report, and its limits are in Section 11.

Cite as

Franco, J. (2026). Treepai g0: A 110M-Parameter Language Model Trained From Scratch on One GPU. Treepai, Inc. https://treepai.com/research/treepai-g0/

@techreport{franco2026treepaig0,
  title       = {Treepai g0: A 110M-Parameter Language Model
                 Trained From Scratch on One GPU},
  author      = {Franco, Jose},
  institution = {Treepai, Inc.},
  year        = {2026},
  month       = oct,
  url         = {https://treepai.com/research/treepai-g0/}
}

Questions about our research: research@treepai.com.

← All research