# Strata Runs Qwen 125B on RTX 4090 at 100T/s

> Published 2026-10-04 · https://www.promptzone.com/arlo_girard/strata-runs-qwen-125b-on-rtx-4090-at-100ts-550d

Strata enables running the **125B parameter Qwen 3.8 Flash Next** model on an **RTX 4090** at **100 tokens per second**. The project first appeared in an [HN thread](https://github.com/Niko1221/Strata) that reached 362 points and 196 comments.

> **Model:** Qwen 3.8 Flash Next | **Parameters:** 125B | **Speed:** 100 tokens/s  
> **Hardware:** RTX 4090 | **License:** Apache 2.0

## What Strata Does

Strata is a set of inference optimizations that quantize and schedule the 125B model for single-GPU consumer cards. It combines 4-bit weight compression with custom kernel scheduling to keep the full model in 24 GB VRAM while sustaining high throughput.

The approach avoids model sharding across multiple GPUs. All computation stays on one RTX 4090.

## Measured Performance

Early reports list consistent **100 tokens per second** on the 4090 for 4k context lengths. Memory footprint stays under 22 GB after quantization.

| Feature          | Strata + Qwen 125B | vLLM (typical) | llama.cpp (4-bit) |
|------------------|--------------------|----------------|-------------------|
| Tokens/s (4090)  | 100                | 35-45          | 25-30             |
| VRAM usage       | 22 GB              | 28+ GB         | 18 GB             |
| Context length   | 4k                 | 4k             | 4k                |
| Multi-GPU needed | No                 | Often          | No                |

## How to Try It

Clone the repository and follow the provided Docker setup. The repo includes a one-line install script that pulls the quantized weights and launches an OpenAI-compatible server.

Users report the full pipeline completes in under ten minutes on a fresh Ubuntu 22.04 install with CUDA 12.4. No additional model conversion steps are required.

## Tradeoffs

Strata delivers speed but requires the exact 4090 configuration tested. Performance drops on 3090 or 4080 cards. Quantization also introduces minor accuracy loss on long-context reasoning tasks compared with 8-bit or FP16 baselines.

The project currently lacks Windows support and multi-user serving features found in production frameworks.

## Alternatives

vLLM and TensorRT-LLM remain stronger for multi-GPU clusters or very long contexts. llama.cpp offers broader hardware compatibility at lower speed. Strata sits between these options for single-GPU, high-throughput use cases.

## Who Benefits

Developers running local agents or chat interfaces on one high-end consumer card gain the most. Teams needing maximum accuracy or serving dozens of concurrent users should evaluate vLLM or hosted APIs instead.

## Verdict

Strata demonstrates that a 125B model can run at interactive speeds on a single 4090 when aggressive quantization and scheduling are applied together. The approach narrows the gap between consumer hardware and large-model inference without requiring cloud resources.

The project is still early, yet the reported numbers already make it worth testing for single-GPU workflows.