GRPO increases the likelihood of samples that are better than average, not just the single best, and decreases that of samples that are worse than average. This method doesn't even involve an explicit likelihood, so it's a completely different mechanism.
A comparison with minibatch optimal transport is in appendix A.2 of the paper.
Without anchoring on the group average it's hardly Group-Relative and without likelihood there's no Policy; that doesn't leave much of the original GRPO unchanged...
Do you have a particular paper in mind that does likelihood-free best-of-K but just calls it GRPO?
Any of the Best of N papers that exploded in popularity after GRPO.
Likelihood is not fundamental to the spirit of GRPO, any exploratory mechanism would work.
That sequential LLMs have a step-wise probability is convenient but not critical to this approach (where rejection sampling is widely used in diffusion models).
The relation to current RLVR methods I think is interesting, they do discuss it a bit but I would be curious to see more about this as well. Quote from the paper:
Exploration beyond Pretraining. The mode collapse XMs address during pretraining also often
shows up in post-training, where RL fine-tuning is known to sharpen models onto a narrow set of
behaviors [76]. Recent fixes such as pass@krewards [77] and best-of-N-aware fine-tuning [78] can
be seen through our lens as Forward XM, with a verifier standing in for ground truth data. These
fixes act only during post-training, though; pretraining with exploration may yield base models that
capture more modes in the first place, leaving RL more to select among.
Minibatch OT in flow matching also has a very similar mechanism, where samples from a noise distribution are matched to the closest data point.
There is a lot of prior work here that the authors neglect to discuss, which portrays this work as more novel than it actually is.