devim-model · report
Why did language learning at 110M fail to translate into measured capability?
As the ~110M plain Transformer moved from 500M to 1.2B supervised tokens, held-out language modeling improved strongly while the frozen capability measure barely moved.
Problem
Did more natural Turkish training produce broader measured capability, or mainly improve language modeling?
Experiment
The same ~110M plain causal Transformer was evaluated at two preregistered exposure gates. The capability instrument remained frozen during the run.
| Measure | Gate A | Gate B | Change |
|---|---|---|---|
| Supervised tokens | 500,019,597 | 1,200,012,349 | +699,992,752 |
| Held-out CE | 4.0734427 | 3.6260754 | −0.4473673 |
| V1.8 macro | 0.380000 | 0.385625 | +0.5625 pp |
| Competencies above chance | 2/8 | 2/8 | 0 |
Result
Language modeling improved unambiguously. The frozen capability instrument did not register a corresponding broad capability gain. The formal decision was INCONCLUSIVE_AT_B, and Stage C was not opened.
What this did not prove
It did not establish a hard 110M capacity ceiling, a Transformer limitation, or the necessity of a new architecture. It instead made data, response interface and evaluator quality separate scientific questions.
Decision
Rather than treating scale as an automatic fix, we moved to bounded FORM, CONTENT and evaluator diagnostics. That split shaped the later QA probes, FORM Rescue and matched-scale experiments.
- Version
- v1
- Evidence
- 2
- Source
- Git · en/reports/110m-language-learning-capability-dissociation.md
DEVİM publications preserve revision history and distinguish public evidence from internal work.