DEVİMDEVİM
← Back

devim-model · report

336M Gate A: did the larger model establish a capacity effect?

The 336M model improved several readouts, but it did not clear the preregistered capacity-effect threshold. Gate B was opened for additional evidence.

Status ValidatedPublished 1 min read168 wordsv1
ListenNot supported in this browser

Why run this experiment?

We did not interpret the 110M capability plateau as "the model is too small" by default. Capacity was isolated and tested with a matched plain-Transformer probe.

Matched Gate A

Measurematched 110Mmatched 336M
Breadth macro0.18750.3635417
FORM EOS0.781250.8541667
FORM repetition0.218750.1458333
V1.8 macro0.3993750.391875
V1.8 competencies above chance32
Positive-control macro0.56473210.7455357

The breadth delta was +0.1760417. The preregistered minimum was +0.20.

Decision

The formal decision was GATE_A_NOT_ESTABLISHED_CONTINUE_GATE_B. capacity_effect_supported remained false.

The larger model improved FORM and positive controls substantially, but that was not enough to establish a general capacity advantage under the frozen gate.

Why it matters

Scale is treated as a testable hypothesis rather than a belief. A larger model may improve some behaviors and still fail a preregistered capability gate.

PUBLICATION RECORDdevim-336m-gate-a
Version
v1
Evidence
1
Source
Git · en/reports/336m-gate-a-matched-scale-probe.md

DEVİM publications preserve revision history and distinguish public evidence from internal work.

336Mscalecapacitymatched experiment