Treepai g0 is our first model and the baseline for our later work. This card is taken from Appendix A of the research report.
| Model | Treepai g0, a 110.1M-parameter decoder-only transformer. Base (pretrained) model only. |
| Developer | Treepai, Inc., October 2026. |
| Availability | Research model. Weights, tokenizer and training code are not released. |
| Intended use | Research at Treepai on training and serving efficiency; a baseline for later models. |
| Out of scope | Any user-facing use, factual question answering, advice of any kind, and languages other than English. |
| Training data | 2.47B tokens from an English FineWeb-Edu set (educational score of 3 or higher; 117 documents that overlapped with benchmarks removed); 63% of the set seen. |
| Evaluation | Nine zero-shot capability benchmarks, held-out bits per byte, and the CrowS-Pairs bias probe (Section 7 of the report). |
| Known risks | States false facts fluently; may reproduce stereotypes; not checked for memorized personal information. |
| Compute | 4.67 GPU-hours on one RTX 4090; at most about 2.1 kWh at the GPU. |
Results in brief
- 37.3% on five core benchmarks, against 38.6% for GPT-2 small and 39.0% for Pythia-160M. With our improved training recipe it reaches 38.9%, level with both.
- Strongest on science questions and weakest on story-level reading (19.4% on LAMBADA against 32.6% for GPT-2 small).
- SmolLM2-135M, trained on about 2 trillion tokens, is well ahead on every benchmark and is our long-term target.
Full results, methods and limits: the report page and the PDF.