Enderchef·inGeneral·3 days agoWhy SFT? And why one distill?Why is every new distill SFT? SFT can't teach as much new knowledge Why not RL? And why does everyone dump old agentic datasets while training on new model ones(e.g. having a Opus…#question1000
Glint Researchmoderator·inGeneral·11 days agoMost wanted small model?What small model would you MOST want to see? It doesn't have to be a LLM, it could be a OCR model, object recognition, anything. What would it be?#question#ideas#hitherehello1020