devim-model · experiment
FORM Rescue: can response behavior improve at the same scale?
Without increasing parameter count, counterbalanced response-register supervision produced large improvements in EOS and repetition behavior.
Question
Was the original 110M model's broken speaking behavior really a capacity problem, or had it simply never learned the response register?
Intervention
A FORM curriculum balanced the same semantic material across different response surfaces. The goal was not to claim reasoning, but to isolate termination, repetition control and usable answer form.
Result
On a 96-item treatment evaluation:
| Measure | Result |
|---|---|
| EOS rate | 0.8541667 |
| Repetition rate | 0.1458333 |
| Empty rate | 0 |
| Mean generated words | 14.99 |
The broader historical V1.8 result record is being independently re-verified; this page therefore publishes only the FORM measurements directly supported by the trained evaluation artifact.
What we learned
The original speaking/interface failure could not be explained by parameter count alone. The same scale learned far more usable response behavior when the supervision register changed.
Boundary
This does not mean CONTENT or reasoning was solved. Better FORM does not by itself show correct content selection or transfer to new relations.
- Version
- v1
- Evidence
- 1
- Source
- Git · en/experiments/form-rescue.md
DEVİM publications preserve revision history and distinguish public evidence from internal work.