Skip to content
r/LocalLLaMA · Communities

README_EN.md · openpangu/openPangu-2.0-Flash at main

1. Introduction openPangu-2.0-Flash is an MoE model trained on Ascend. The model has 92B total parameters and 6B activated parameters. Its context length is 512k. The total pretraining data contains 34T tokens. During Post-training, openPangu-2.0-Flash is trained through unified SFT with slow and fast thinking capabili