Testing whether Lojban's regular, unambiguous grammar gives small language models an advantage over English
These 570K-parameter models run in your browser. Give them a bAbI-style prompt and see what they generate.
Lojban is a constructed language designed for logical, unambiguous communication. Every sentence has exactly one parse tree — no garden paths, no structural ambiguity.
"Time flies like an arrow."
3+ valid parses. Are we timing flies? Do time-flies enjoy arrows?
"lo temci cu vofli tai lo bagre"
Exactly 1 parse. "Time flies in-the-manner-of an arrow."
Every sentence has exactly one syntactic analysis. The grammar is a formal PEG, machine-parseable.
Word categories determined by shape: CVCCV = verb, CCVCV = verb, CVC+V = name. No irregular forms.
Grammatical roles marked by particles, not word order. lo = article, cu = predicate marker, pu = past tense.
Type text to see how each language gets tokenized with BPE (vocab=1024)
Five iterations of hypothesis → confound → fix → repeat
Lojban achieves 100% grammaticality at every model size. English improves from 73% to 99% but never reaches perfection.
Bits-per-character on held-out text. Lower is better. Lojban consistently 15-35% lower across all experiments.
Validation BPC over training steps (V4 medium, seed 42). Lojban converges faster and to a lower minimum. Both overfit after ~1K steps.
English's bAbI advantage was an artifact of training duration, not reasoning. As we fixed confounds, the gap shrank from 26pp to 1.3pp.
Generated text from V4 medium models (570K params, BPE tokenization, seed 42)