Home › Migrating from Llama 3.3 to GPT-OSS 120B

Llama 3.3 70B Is Being Retired on Groq. Here's How I Migrated.

Groq is decommissioning llama-3.3-70b-versatile on August 16, 2026. I run a free AI desktop app on that model. This is what the migration actually involved — including the benchmarks, the limits the docs gloss over, and the mistakes that cost me time.

By Bhavik Thakkar · Creator of BEBO the PET · August 2026

The email nobody wants

Groq sent a deprecation notice: llama-3.3-70b-versatile would be decommissioned on August 16, 2026, with openai/gpt-oss-120b and qwen/qwen3.6-27b as the recommended replacements. This applies to free and developer-tier usage; committed-spend enterprise contracts are unaffected.

I maintain BEBO the PET, a free open-source AI assistant that lives on your Windows desktop. Llama 3.3 70B was its entire brain. Every summarise, every email draft, every grammar fix ran through it. After August 16, every one of those buttons would return an error for every user who had ever installed the app.

The important reframe: this is not a bug you introduced, and it is not something you can patch around. An upstream provider is removing a model that every app built on it depends on. Your only real choice is which replacement you move to, and how quickly.

Step 1: Benchmark before you believe anything

The recommended replacement being "recommended" doesn't tell you whether it's actually good for your workload. Before changing a single line of production code, I ran the same task through all three models with identical parameters — same system prompt, same temperature, same max_tokens.

The test was a deliberately broken sentence for grammar correction, which is BEBO's most latency-sensitive task:

// Same params across all three: temperature 0.05, top_p 0.9, max_tokens 500
const body = {
  model,
  temperature: 0.05,
  max_tokens: 500,
  messages: [
    { role: 'system', content: SYSTEM_PROMPT },
    { role: 'user',   content: BROKEN_SENTENCE }
  ]
};

Measured results, using each response's own completion_time from the API rather than wall-clock time:

ModelThroughputOutput qualityVerdict
openai/gpt-oss-120b~481 tok/sCorrect, natural phrasingChosen — primary
openai/gpt-oss-20b~971 tok/sCorrect, slightly clumsierChosen — fallback
llama-3.3-70b-versatile~428 tok/sCorrectBeing retired

The surprise: the replacement was faster than the model being retired — roughly 12% higher throughput on the same task. I had assumed a forced migration meant accepting a downgrade. It didn't. Benchmark before you assume.

Step 2: The free-tier limits nobody mentions

This is where I found a genuine problem — not with the migration, but with claims I had been making for months.

My README, website and launch materials all said BEBO gave you 14,400 free requests per day. That number came from multiplying the requests-per-minute limit by the minutes in a day. It was never real. Groq enforces a separate, much lower daily cap.

The actual documented free-tier limits:

Limitgpt-oss-120bgpt-oss-20bllama-3.3-70b
Requests / minute303030
Requests / day1,0001,0001,000
Tokens / minute8,0008,00012,000
Tokens / day200,000200,000100,000

So the honest figure was 1,000 requests per day — about 40 AI tasks every hour, all day, which is still far more than any realistic desktop user needs. But it wasn't 14,400, and it never had been.

If you take one thing from this post: a forced migration is a good moment to audit every number you have published. I found an inaccurate claim repeated across 25+ files. It wasn't caused by the deprecation — the deprecation just made me look.

Step 3: Build a fallback chain, not a swap

The naive fix is to change one string and ship. I did something slightly more durable: a chain of models tried in order, so a busy or rate-limited primary doesn't take the app down.

// One constant. The next migration is a one-line change.
const GROQ_MODELS = [
  'openai/gpt-oss-120b',  // primary — best quality
  'openai/gpt-oss-20b'    // fallback — faster, separate quota
];

for (const model of GROQ_MODELS) {
  const res = await callGroq(model, payload);
  if (res.ok) return res;
  // 429 or 5xx — try the next model instead of failing
}

Two benefits beyond resilience. The fallback has its own quota, so hitting a limit on one model doesn't stop the app. And critically, the model name now lives in exactly one place. When Groq retires GPT-OSS someday, the code change is a single line.

The mistake that cost me the most time

Before this refactor, the model name was hardcoded in two files — the AI service and the API-key validation routine. I changed one and not the other. Key validation silently kept calling a model that was about to disappear. It would have kept working right up until it didn't.

Grep for your model string across the entire repository, not just the file you think owns it.

Step 4: Migrate the users, not just the code

Shipping the new version fixes nothing for people who already installed the old one. Desktop apps don't refresh like websites do — an installed binary keeps running the retired model until the user actively updates.

What actually reaches existing users:

A CI trap worth knowing: my build pipeline failed twice after the version bump. First because package-lock.json still carried the old version and npm ci refuses to run on a mismatch. Then because electron-builder auto-enables publishing when it detects CI and fails without a token — fixed with --publish never, since releases are published deliberately, not by the build.

What I'd do differently

  1. Keep the model name in one constant from day one. Hardcoding it twice is how you ship a half-migration.
  2. Ship the update checker before you need it. I built BEBO's in-app update notice during this migration. It would have been far more useful shipped a version earlier, already sitting in every install.
  3. Treat published numbers as code. Marketing claims scattered across a README, a website, a deck and image assets drift out of sync silently. Mine did.
  4. Benchmark the replacement before assuming a downgrade. Mine was faster.

The short version

A provider retiring your model is genuinely disruptive, but it is also finite and fixable. Benchmark the replacements with your own workload, verify the limits rather than deriving them, put the model name behind one constant with a fallback, and remember that migrating the code is only half the job — the installed users are the other half.

BEBO now runs on GPT-OSS 120B with an automatic GPT-OSS 20B fallback. It's faster than before, still completely free, and the next migration will be a one-line change.

Further reading

I've also written about BEBO and this migration on dev.to:

BEBO the PET robot

Try BEBO the PET

A free, open-source AI assistant that lives on your Windows desktop. No subscription, no browser tabs.

⇓ Download Free

Explore more