deepseek-r1-rl-without-sft
IN premise — entries/2026/06/21/wiki-Reinforcement_learning-chunk-4.md
Created 2026-06-21T09:55:53+00:00
DeepSeek-R1 uses large-scale RL without supervised fine-tuning (SFT) as a preliminary step and achieves reasoning performance comparable to OpenAI-o1-1217
Dependents
These beliefs depend on this one:
- IN deepseek-validates-paradigm-taxonomy-dissolution-in-rl — DeepSeek-R1's achievement of competitive reasoning performance through large-scale RL without supervised fine-tuning further validates paradigm taxonomy dissolution — a traditionally supervised task (reasoning) solved through RL alone, without the supervised intermediate step that the standard LLM pipeline assumes, demonstrating that paradigm boundaries dissolve not only in training pipelines but in task requirements.