A mixture-of-experts model that punches above its weight by activating only 3 billion parameters per forward pass despite having 30 billion total parameters. This selective activation makes it computationally lean while retaining a broad knowledge base. Its 262K token context window is notably large, letting it handle lengthy documents without truncation.