DeepSeek-MoE

核心与官方官方模型与基建MIT官方停更
Model Download | Evaluation Results | Quick Start | License | Citation

功能特性

  • Comparison with LLaMA2 7B on our internal benchmarks. With only 39.6% of computations, DeepSeekMoE 16B outperforms LLaMA2 7B on the majority of benchmarks.

项目预览

项目简介

DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

← 返回 核心与官方 列表