Model Download | Evaluation Results | Quick Start | License | Citation
功能特性
- Comparison with LLaMA2 7B on our internal benchmarks. With only 39.6% of computations, DeepSeekMoE 16B outperforms LLaMA2 7B on the majority of benchmarks.
项目简介
DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models


