DEVİMDEVİM
← Back

devim-model · report

Why did language learning at 110M fail to translate into measured capability?

As the ~110M plain Transformer moved from 500M to 1.2B supervised tokens, held-out language modeling improved strongly while the frozen capability measure barely moved.

Status ValidatedPublished 1 min read190 wordsv1
ListenNot supported in this browser

Problem

Did more natural Turkish training produce broader measured capability, or mainly improve language modeling?

Experiment

The same ~110M plain causal Transformer was evaluated at two preregistered exposure gates. The capability instrument remained frozen during the run.

MeasureGate AGate BChange
Supervised tokens500,019,5971,200,012,349+699,992,752
Held-out CE4.07344273.6260754−0.4473673
V1.8 macro0.3800000.385625+0.5625 pp
Competencies above chance2/82/80

Result

Language modeling improved unambiguously. The frozen capability instrument did not register a corresponding broad capability gain. The formal decision was INCONCLUSIVE_AT_B, and Stage C was not opened.

What this did not prove

It did not establish a hard 110M capacity ceiling, a Transformer limitation, or the necessity of a new architecture. It instead made data, response interface and evaluator quality separate scientific questions.

Decision

Rather than treating scale as an automatic fix, we moved to bounded FORM, CONTENT and evaluator diagnostics. That split shaped the later QA probes, FORM Rescue and matched-scale experiments.

PUBLICATION RECORDdevim-110m-capability-dissociation
Version
v1
Evidence
2
Source
Git · en/reports/110m-language-learning-capability-dissociation.md

DEVİM publications preserve revision history and distinguish public evidence from internal work.

110Mnegative resultevaluationTurkish