Can Grammar Help
Tiny Models Think?

Testing whether Lojban's regular, unambiguous grammar gives small language models an advantage over English

100% Lojban grammar
at 67K params
15-35% Lower prediction
error (BPC)
5 Experiment
iterations
Try the models yourself ↓

Try It Yourself

These 570K-parameter models run in your browser. Give them a bAbI-style prompt and see what they generate.

English

Loading model...
0.8
128
Generated text will appear here...

Lojban

Loading model...
0.8
128
Generated text will appear here...

What is Lojban?

Lojban is a constructed language designed for logical, unambiguous communication. Every sentence has exactly one parse tree — no garden paths, no structural ambiguity.

English

"Time flies like an arrow."

3+ valid parses. Are we timing flies? Do time-flies enjoy arrows?

Lojban

"lo temci cu vofli tai lo bagre"

Exactly 1 parse. "Time flies in-the-manner-of an arrow."

Unambiguous Parse

Every sentence has exactly one syntactic analysis. The grammar is a formal PEG, machine-parseable.

Regular Morphology

Word categories determined by shape: CVCCV = verb, CCVCV = verb, CVC+V = name. No irregular forms.

Explicit Structure

Grammatical roles marked by particles, not word order. lo = article, cu = predicate marker, pu = past tense.

Try the BPE Tokenizer

Type text to see how each language gets tokenized with BPE (vocab=1024)

The Journey

Five iterations of hypothesis → confound → fix → repeat

Key Findings

Grammar at Every Scale

Lojban achieves 100% grammaticality at every model size. English improves from 73% to 99% but never reaches perfection.

Prediction Quality (Test BPC)

Bits-per-character on held-out text. Lower is better. Lojban consistently 15-35% lower across all experiments.

Training Dynamics

Validation BPC over training steps (V4 medium, seed 42). Lojban converges faster and to a lower minimum. Both overfit after ~1K steps.

The bAbI Confound Story

English's bAbI advantage was an artifact of training duration, not reasoning. As we fixed confounds, the gap shrank from 26pp to 1.3pp.

Sample Comparisons

Generated text from V4 medium models (570K params, BPE tokenization, seed 42)

Methodology

Model Architectures
Training Details
  • Architecture: Decoder-only Transformer (GPT-style)
  • Optimizer: AdamW (lr=3e-4, weight_decay=0.1, grad_clip=1.0)
  • Training data: 4 parallel books (~407K chars/language) + 20 bAbI tasks (~2.4M chars)
  • Test data: Metamorphosis (held-out, ~120K chars) + bAbI test sets (seen/unseen vocab)
  • Seeds: 42, 137, 2024 (3 runs per condition)
  • V4 specifics: BPE tokenization (vocab=1024), 10K fixed steps, checkpoint by best val BPC
Key Lessons Learned
  1. Early stopping conflates entropy with capability. Character-prediction convergence speed is orthogonal to reasoning ability. V3→V3.1 showed this dramatically.
  2. Character-level tokenization penalizes multi-word expressions. Lojban's 2-3 token locations vs English's single tokens dominated at tiny scale.
  3. Checkpoint selection must align with evaluation objective. Selecting by narrative val loss then evaluating bAbI accuracy is misaligned.
  4. Neither language shows reasoning below ~1M params. bAbI tasks require more capacity than grammar or pattern matching.