A practical map of how GRPO-family methods estimate advantages, constrain policy updates, and balance reward improvement with continued exploration.
When a detector decides what gets investigated, it shapes the labels used to train and evaluate its successor.