Why is every new distill SFT? SFT can't teach as much new knowledge Why not RL? And why does everyone dump old agentic datasets while training on new model ones(e.g. having a Opus 4.8 and Fable 5 dataset, training on the Fable 5 dataset and not the Opus 4.8) when the slightly older model ones have good data too, and the prompts are unique? Why not make a model on several different datasets?