devim-model · experiment
What did a failed QA experiment teach us?
Answering behavior improved partially, but held-out CE regressed and a later control showed that much of the task could be solved by trivial copying.
Problem
The 110M parent had learned natural text modeling but struggled to answer directly, stop, and control repetition. We wanted to separate missing capability from a missing response interface.
Hypothesis
A small source-grounded QA dose might teach answer form and termination. If the instrument was clean, we could then examine whether content behavior improved as well.
v0.1 result
| Measure | Parent | Child |
|---|---|---|
| Exact span | 0.0000 | 0.0150 |
| EOS | 0.0725 | 0.5150 |
| Repetition | 0.8725 | 0.3800 |
| V1.8 macro | 0.385625 | 0.411875 |
| Held-out CE | 3.62607545 | 4.17608298 |
Response form improved materially, but the success bar was not reached and natural language modeling regressed. The child was rejected as a future parent.
v0.2 control: the copy problem
| Family | Mechanical copy baseline | Neural child | Margin |
|---|---|---|---|
| F1 | 0.78 | 0.24 | −54 pp |
| F2 | 0.94 | 0.52 | −42 pp |
A substantial part of the task was solvable by trivial copying. An extractive score could therefore not be treated as evidence of relational understanding by itself.
Decision
We did not hide the failure. The evaluator/task confound became the result. Negative controls and copy-solvability checks became mandatory parts of later evaluator design.
- Version
- v1
- Evidence
- 2
- Source
- Git · en/experiments/qa-interface-and-copy-confound.md
DEVİM publications preserve revision history and distinguish public evidence from internal work.