devim-model · report
336M Gate A: did the larger model establish a capacity effect?
The 336M model improved several readouts, but it did not clear the preregistered capacity-effect threshold. Gate B was opened for additional evidence.
Why run this experiment?
We did not interpret the 110M capability plateau as "the model is too small" by default. Capacity was isolated and tested with a matched plain-Transformer probe.
Matched Gate A
| Measure | matched 110M | matched 336M |
|---|---|---|
| Breadth macro | 0.1875 | 0.3635417 |
| FORM EOS | 0.78125 | 0.8541667 |
| FORM repetition | 0.21875 | 0.1458333 |
| V1.8 macro | 0.399375 | 0.391875 |
| V1.8 competencies above chance | 3 | 2 |
| Positive-control macro | 0.5647321 | 0.7455357 |
The breadth delta was +0.1760417. The preregistered minimum was +0.20.
Decision
The formal decision was GATE_A_NOT_ESTABLISHED_CONTINUE_GATE_B. capacity_effect_supported remained false.
The larger model improved FORM and positive controls substantially, but that was not enough to establish a general capacity advantage under the frozen gate.
Why it matters
Scale is treated as a testable hypothesis rather than a belief. A larger model may improve some behaviors and still fail a preregistered capability gate.
- Version
- v1
- Evidence
- 1
- Source
- Git · en/reports/336m-gate-a-matched-scale-probe.md
DEVİM publications preserve revision history and distinguish public evidence from internal work.