JundeWu2025-01-27 21:15:06我之前说OpenAI>Deepseek,所以追赶OpenAI的临门一脚到底差在哪里?先说结论,我认为是超高质量的RLHF (Reinforcement Learning from Human Feedback),也就是人类反馈 Deepseek-R1这次的训练,仅利用了rule-based outcome reward,也就是数学题答案对错/测代码能不能跑通,训练出来了超强的逻辑,在math/c#OpenAI#DeepSeek#RLHF#Reinforcement Learning from Human Feedback#rule-based outcome reward#逻辑能力#math/c