A lean, fast-moving model that punches above its weight by activating only 3B parameters out of 30B total — a mixture-of-experts design that keeps inference costs low without sacrificing much capability. The FP8 quantization further trims memory footprint, making it practical to run on modest hardware. It handles a remarkably long context window of 262K tokens, so it can work through large documents without losing the thread.