You're entirely right, I reworked everything today, added IFEval and GPQA to the benchmark, and GAIA will come after for the autonomy bench.
At the moment, the first step is training the model to train itself alone, it's in a constant loop where he train itself, launch a script that applies this new learning on itself with a LoRA made out of the iteration he just did, bench itself, check if the result are better or worse, when an iteration is worse he don't keep it, when it's a better one he keep it, and the loop start again.
When I will be able to get 100% on the two first column, then I will add GAIA bench and start to train it with the actual knowledge it took by himself.
But right now, I need it to be consistent for what I want it to do. The model is still running, but since I added more benchmark and modified part of the code in the loop this morning, it take more time, probably will do one iteration per day, I will post the new result later and dismiss the first batch.
Your help is valuable and I will re-check everything before doing anything when this loop close!
Edit: Autonomy eval was 12 episode, 6 in french and 6 in english, modified that today for 50 in french and 50 in english with web caching to avoid modification post-train. I also modified the number episodes he need for each train: from 40 to 300, since the little test I did was successful.
Edit 2 : A correction first, my post was wrong. Learned facts were not part of the gate (oops). Also, in that early table the fact questions were regenerated at every export, so 0.12 -> 0.05 compared different question sets. On the same questions, base scored about 0.04โ0.05. So it wasn't a measured regression, but it wasn't a valid measurement either.
Since then, the evaluation protocol has been frozen:
- The base weights go through the exact same pipeline as every candidate (same merge path, same Q8_0 conversion, same Modelfile), so only training differs.
- Paired comparisons against that frozen base, never against the previous round, item by item.
- Autonomy on 100 fixed stimuli, with a fixed seed per episode and a cached web, so every model sees the same pages for the same queries.
- A frozen 400-question facts panel, and a hash manifest so results from different protocols are never compared.
After your comment, the gate itself changed:
- Held-out facts are now gated.
- Autonomy is compared to the best model ever promoted, not the previous one, to block exactly the ratchet you describe.
- Any missing measurement now means rejection, instead of being skipped.
On noise: an A/A run starts tonight. The frozen base is re-run with different seeds against its own official results. It must pass its own gate, and it gives the real noise floor per benchmark.
On train facts: agreed. So far the loop taught the model how to act, not what it read. That's the next phase, and it will be measured on the same frozen panel.
Will post new bench when actual progress happen. Time to get out the big guns kek.