Back to the blog
Local AIBunLLMOpen Source

Qwen 3.8 125B on an RTX 4090: what changes for Bun users

October 05, 2026·6 min read·Diego Horvatti

A repository called Strata showed up on GitHub with a bold promise. It runs Qwen 3.8 Flash Next, a 125 billion parameter model, on an RTX 4090 at 100 tokens per second. That's a gaming card with 24 GB of VRAM. If this holds up, a model that needed a datacenter server until yesterday now fits under your desk. In this post I want to separate the physics from the marketing. Then I'll show how to wire Qwen 3.8 Flash Next into a Bun backend without any drama.

What Strata promises, exactly

Strata is an inference runtime. It doesn't train anything or create a new model. Its job is to take a big model and make it run on hardware that, in theory, shouldn't handle it.

The headline has three numbers: 125B parameters, one RTX 4090 and 100 tokens per second. Each one alone is reasonable. All three together catch your eye.

For comparison, 100 tokens per second is faster than you can read. At that speed text looks "dumped" on the screen, not typed. Plenty of paid APIs deliver less than that at peak hours.

How does a 125B model fit in 24 GB of VRAM?

Let's do some back-of-the-envelope math.

  • In FP16, each parameter takes 2 bytes. 125 billion × 2 = 250 GB.
  • Quantized to 4 bits, it drops to about 62 GB.
  • The 4090 has 24 GB.

It doesn't fit. Not even close. So there's a trick, and the likely trick is in the name. "Flash" usually points to a sparse model, a Mixture of Experts (MoE). In this format, the model has 125B parameters in total, but only a fraction of them works on each token. A router picks half a dozen "experts" and the rest sits idle.

That changes everything. If only a small part of the model fires per token, you don't need the whole model on the GPU. You need the right experts at the right time. The rest can live in system RAM.

And that's where the real bottleneck shows up. The 4090's memory delivers close to 1 TB/s. The PCIe 4.0 bus between RAM and the card delivers about 25 GB/s in practice. That's a 40x gap. Every expert that has to come from RAM mid-generation is an expensive toll.

So the 100 tokens per second depends less on the GPU and more on one question: how often was the expert the model asked for already in VRAM? If Strata hits that cache almost every time, the number is plausible. If your prompt falls outside the pattern, speed falls off a cliff.

A big model on a small card isn't magic. It's good caching.

What to measure before you believe 100 tokens/s

A README number is a README number until you measure it on your machine. Before you reshuffle your stack, check these points:

  • Prefill vs. decode. Generating token by token (decode) is one thing. Processing a 20,000 token prompt before answering (prefill) is another. Many benchmarks only show decode.
  • Context size. With 500 tokens of conversation, everything flies. With your whole project file pasted into the prompt, the attention cache also fights for VRAM.
  • Batch 1. Almost every homemade number is for a single user. Two requests at once can split the speed, or worse.
  • System RAM. If the leftover experts live in RAM, you need 64 GB or more. The 4090 alone won't cut it.
  • Quantization quality. 4 bits on an MoE model can turn out great, or it can botch basic SQL. Test with your own tasks, not someone else's benchmark.

Tokens per second in a README is like fuel economy in a car ad. It's true, just downhill with a tailwind.

How to plug a local model into a Bun backend

Now the practical part. Most local inference runtimes expose an endpoint compatible with the OpenAI API (/v1/chat/completions). Check Strata's README to see if it does this out of the box. If it doesn't, any thin server in front of it will do. The port and model name below are examples, so adjust them to your setup.

With Bun you don't need any SDK. Native fetch already handles streaming:

// chat-local.ts
const res = await fetch("http://localhost:8080/v1/chat/completions", {
  method: "POST",
  headers: { "Content-Type": "application/json" },
  body: JSON.stringify({
    model: "qwen3.8-flash-next",
    stream: true,
    messages: [{ role: "user", content: "Explain: ERROR: deadlock detected" }],
  }),
});

const decoder = new TextDecoder();
let buffer = "";
let chunks = 0;
const start = performance.now();

for await (const part of res.body!) {
  buffer += decoder.decode(part, { stream: true });
  const lines = buffer.split("\n");
  buffer = lines.pop() ?? "";
  for (const line of lines) {
    if (!line.startsWith("data: ") || line.includes("[DONE]")) continue;
    const delta = JSON.parse(line.slice(6)).choices[0]?.delta?.content;
    if (delta) {
      chunks++;
      process.stdout.write(delta);
    }
  }
}

const secs = (performance.now() - start) / 1000;
console.log(`\n~${(chunks / secs).toFixed(1)} tokens/s`);

Run it with bun chat-local.ts and you're done. As a bonus, you get your own tokens per second measurement. The math is approximate, since most servers send one token per chunk. Still, it's good enough to compare against the README's promise.

Two details here are worth their weight in gold. First, the timing includes prefill, so test with a short prompt and a long one and compare. Second, since the endpoint follows the OpenAI format, switching between a local model and a paid API becomes an environment variable:

const BASE_URL = Bun.env.LLM_BASE_URL ?? "http://localhost:8080/v1";

In dev you point it at your 4090. In production, at the provider. The code doesn't change.

Is it worth swapping a paid API for a 4090?

It depends on what you do with AI day to day. An RTX 4090 costs more than a lot of laptops. It also comes with a power bill and a room that turns into a sauna.

Where it really makes sense:

  • Sensitive data. Client code, contracts, personal data. Running locally settles a whole GDPR conversation before it even starts.
  • Batch jobs. Classifying 50,000 tickets, writing product descriptions, summarizing logs. No rate limits and no surprise invoice at the end of the month.
  • Internal dev tools. Reviewing diffs, explaining Postgres errors, generating test fixtures. One person using it, which is exactly the scenario where the batch 1 benchmark holds.

Where it doesn't:

  • A product with many concurrent users. A home card is not an inference cluster. Ten people at once and the 100 tokens/s becomes a memory.
  • When you need the best model on the market. A quantized 125B is good. It's not the top, and pretending it is will only cost you rework.

My take on Strata

The number that matters in this story isn't the 100. It's the 125. Running a model at this scale on consumer hardware, even at 40 tokens per second, already changes who can experiment with AI without asking anyone for a budget. The rest is optimization, and optimization gets better every month.

Now, my strong opinion: don't build architecture on top of a README benchmark. Measure on your machine, with your prompt. And put the model URL in an environment variable from day one. With Bun that costs ten lines. It also leaves you free to switch models when the next miracle repo shows up, which should be sometime next week.

If you want to see how I wire this kind of thing into real projects, take a look at my projects.

LinkedIn summary

A 125 billion parameter model running on a gaming card. Sounds like a joke, but that's what Strata promises.

The math doesn't add up. Even quantized, Qwen 3.8 needs about 62 GB, and the RTX 4090 has 24. The trick is MoE: only a few experts work on each token, and the rest stays in system RAM.

So the 100 tokens/s depends less on the GPU and more on how often the cache hits. A big model on a small card isn't magic. It's good caching.

My rule: don't build architecture on top of a README benchmark. Measure on your machine, with your prompt.

And put the model URL in an environment variable from day one. With Bun and native fetch it's about ten lines, and switching between a local model and a paid API becomes a config change.

I wrote the full walkthrough with the complete code on the blog. If you run local AI in your backend, tell me how you're measuring it.

#AI #Bun #LLM #Qwen #SoftwareDevelopment