Model card

Treepai g0

Research model · October 2026

Treepai g0 is our first model and the baseline for our later work. This card is taken from Appendix A of the research report.

ModelTreepai g0, a 110.1M-parameter decoder-only transformer. Base (pretrained) model only.
DeveloperTreepai, Inc., October 2026.
AvailabilityResearch model. Weights, tokenizer and training code are not released.
Intended useResearch at Treepai on training and serving efficiency; a baseline for later models.
Out of scopeAny user-facing use, factual question answering, advice of any kind, and languages other than English.
Training data2.47B tokens from an English FineWeb-Edu set (educational score of 3 or higher; 117 documents that overlapped with benchmarks removed); 63% of the set seen.
EvaluationNine zero-shot capability benchmarks, held-out bits per byte, and the CrowS-Pairs bias probe (Section 7 of the report).
Known risksStates false facts fluently; may reproduce stereotypes; not checked for memorized personal information.
Compute4.67 GPU-hours on one RTX 4090; at most about 2.1 kWh at the GPU.

Results in brief

  • 37.3% on five core benchmarks, against 38.6% for GPT-2 small and 39.0% for Pythia-160M. With our improved training recipe it reaches 38.9%, level with both.
  • Strongest on science questions and weakest on story-level reading (19.4% on LAMBADA against 32.6% for GPT-2 small).
  • SmolLM2-135M, trained on about 2 trillion tokens, is well ahead on every benchmark and is our long-term target.

Full results, methods and limits: the report page and the PDF.

← All models

Research updates

We email when we publish something new. Unsubscribe at any time. Privacy notice.