OpenLearnLM/special-r1-deepseek-qwen3-8b-sped-adaptive-think-noreward Text Generation • 8B • Updated 1 day ago • 181
OpenLearnLM/special-r1-deepseek-qwen3-8b-sped-adaptive-think-noreward Text Generation • 8B • Updated 1 day ago • 181
Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning Paper • 2510.19338 • Published Oct 22, 2025 • 117