A lean, efficient text model that punches above its weight through a mixture-of-experts architecture — only 3 billion parameters are active at a time despite the 30 billion total, keeping inference costs low. It runs in a quantized W4A16 format, trading a small amount of precision for significantly reduced memory footprint. Its standout feature is an unusually large 262K token context window, allowing it to work across very long documents without truncation.