devim-model · research
What went wrong in DEVİM?
We track progress through failed gates, evaluator confounds, overfitting and rejected diagnoses—not only through successful experiments.
Why publish failures?
A research program told only through successful results hides the real learning process. DEVİM keeps wrong diagnoses, confounded evaluators and rejected child models in the record.
1. We could have mistaken lower CE for capability
At 110M, held-out CE improved strongly while the capability measure barely moved. Better language modeling and better performance on a relational instrument are not the same claim.
2. We could have blamed broken speaking on capacity
The original model rarely stopped and repeated heavily. FORM Rescue improved those behaviors at the same parameter scale. A major part of the failure was the supervision register.
3. The QA evaluator could have misled us
Some v0.1 child scores improved. The later v0.2 control showed mechanical copy baselines beating the neural child by 42–54 points. The task rewarded copying more than the capability we intended to measure.
4. We rejected a diagnostic child instead of promoting it
The QA child improved EOS but regressed held-out CE by roughly 0.55 nats. It was rejected as a future parent. One improved metric was not enough for lineage promotion.
5. We did not declare 336M an automatic solution
The larger model improved several readouts but missed the preregistered breadth-delta gate. capacity_effect_supported therefore remained false.
Research principle
Failure is not something to hide; it is evidence used to separate causes. This page will be versioned as new important negative results appear.
- Version
- v1
- Evidence
- 4
- Source
- Git · en/research/what-went-wrong-in-devim.md
DEVİM publications preserve revision history and distinguish public evidence from internal work.