A sparse mixture-of-experts model that punches above its weight by activating only 3 billion parameters per forward pass despite having 30 billion total. This architecture makes it computationally lean at inference time while retaining the capacity of a much larger model. It handles long contexts exceptionally well, supporting up to 262,144 tokens — roughly the length of several novels.