💡Can open-source models verify their own outputs and beat frontier models doing so?
Check out @jackyk02's work on verification scaling: sampling just 5 solutions with DeepSeek V4 Flash + LLM-as-a-Verifier outperforms Claude Fable 5 on Terminal-Bench at 11x lower cost.👇
Scaling self-verification with DeepSeek V4 Flash beats Claude Fable 5 on Terminal-Bench 2.1, while being 11x cheaper 💰
As open-source models become more capable, they can now generate large numbers of high-quality candidate solutions and verify their own outputs at very low





