Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> Deepseek obviously trained on OpenAI outputs

I’ve seen this claim but I don’t know how it could work. Is it really possible to train a new foundational model using just the outputs (not even weights) of another model? Is there any research describing that process? Maybe that explains the low (claimed) costs.



Probably not the whole model, but the first step was "fine tuning" the base model on ~800 chain of thought examples.

Those were probably from OpenAI models. Then they used reinforcement learning to expand the reasoning capabilities.


800k. They say they came from earlier versions of their own models, with a lot of bad examples rejected. They don't seem to say which models they got the "thousands of cold-start" examples from earlier in the process though.


every single model does/did this. Initially fine tuning required the expensive hand labeled outputs for RLHF. Generating your training data from that inherently encodes the learned distributions and improves performance, hence why some models would call themselves chatgpt despite not being openai models.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: