curious where the disagreement lands: the claim i'm least sure of myself is that measurement alone already counts as activation (nothing in the weights changes, so it's a looser sense of the word than usual)
the part i'd defend harder is the eval -> reward one: once a benchmark becomes the thing you train against, its flaws stop being measurement error and start being incentives. if you're pushing back somewhere in there, i'd genuinely like to hear it
yeah its really counterintuitive i think; i.e, getting the right framework and structure for this to work probably isn't trivial, models really hate playing well together. i wonder how their version would fair in real world use.