A thinking-mode specialist that operates as a massive mixture-of-experts architecture — 2.4 trillion total parameters with 95 billion active at any time. It handles long-context tasks gracefully with a 262K token window, making it comfortable with very lengthy documents. As an open-weight model, it can be self-hosted, though the sheer scale of the architecture demands serious infrastructure to run.