@Drewch
Post
Post 1 of 6
I'm realizing from the flywheel thread from @tobi that many people don't realize that small models can actually outperform larger models on a specific task - @MParakhin has been saying this for quite some time. đź§µ
Post 2 of 6
The key is a data flywheel, you need to be constantly identifying incorrect outputs from production distribution, and use high reasoning models + escalate to humans to correct the output, or specifically, give hints to the actual model during training to output correctly.
Post 3 of 6
The above is known as on-policy self distillation. It's a powerful way to get the exact reasoning capabilities you need into the weights, so that at inference, you don't need to do high reasoning.
Post 4 of 6
It's also a much more effective way to improve the system rather than keep adding more and more to your system prompt. It's much nicer working in the continuous space for improvements with gradient descent, rather than in the discrete lexical space.
Post 5 of 6
Surely, there must be some downside to using a smaller model? Yes, there is, but the upsides outweigh the downside. You will lose generalizability, but you will vastly improve quality while reducing latency and cost.
Post 6 of 6
Check out the talk at @icmlconf that @cmazzaanthony and I gave this summer. Stay tuned for @Shopify in December at @NeurIPSConf. https://icml.cc/virtual/2026/75732
Explanation
Andrew McNamara of Shopify is making a fairly specific engineering claim: once an LLM task is narrow, high-volume, and measurable, the best production architecture may not be “call the smartest frontier model with an increasingly elaborate prompt.” Instead, use the frontier model mainly as a teacher and error-corrector, then continuously train a much smaller model on the exact distribution of requests your product actually sees. Shopify says this approach lets its specialized models beat frontier models on particular production tasks while being much faster and cheaper. Tobi Lütke’s post that triggered the thread gave the striking example: a fine-tuned 0.8B-parameter model beating a much larger frontier reasoning model on one very specialized Shopify task. ([Nitter][1])
The important phrase is “production distribution.” It means: don’t train on some generic benchmark or arbitrary synthetic dataset. Collect the actual kinds of inputs users send your deployed system, especially cases where your current model fails. Shopify’s loop is roughly:
user request → model answers → evaluator detects a bad answer → stronger model tries to repair it → sometimes a human supplies the missing insight → repaired trajectory becomes training data → small production model is updated → repeat.
That is the “flywheel.” Production usage continually discovers exactly what the model still cannot do. Shopify describes its version as mining low-scoring conversations, repairing them, and feeding the successful trajectories back into supervised fine-tuning and reinforcement learning. ([Shopify][2])
“On-policy self-distillation” is the densest bit. Ordinary distillation often looks like: take a strong teacher model, have it answer a large preconstructed dataset, and train a smaller student to imitate those answers. Here, “on-policy” means the examples are generated from situations encountered by the model/system you are actually deploying—especially its own current failures. “Self-distillation” is being used somewhat broadly: stronger reasoning/model machinery produces a corrected trajectory that teaches the cheaper model what it should have done. Crucially, the teacher can provide more than the final answer: it can supply hints or intermediate guidance that makes the desired behavior learnable.
“In the weights” versus “in the prompt” is the other central distinction. Suppose the model repeatedly makes mistake X. You can append another rule to a system prompt: “When X occurs, remember to do Y.” Do this 500 times and you get a giant fragile instruction manual. Alternatively, generate examples demonstrating the correct behavior and train on them. Gradient descent then changes millions or billions of parameters so that doing Y becomes part of the model’s learned behavior. That is what McNamara means by moving from “discrete lexical space”—explicit English instructions—to a “continuous space” of learned weights. Shopify explicitly frames its continual-learning system this way: production knowledge gets compressed into parameters rather than accumulating indefinitely as prompt text and harness code. ([Shopify][2])
Why can the tiny model actually beat the giant one? Because parameter count buys breadth. A frontier model has to know how to write poetry, reason about physics, speak dozens of languages, program in obscure frameworks, etc. A 0.8B model can devote effectively all of its limited capacity to one narrow mapping. This is not merely theoretical; research also finds that specialized smaller models can equal or exceed untuned larger general-purpose models on sufficiently constrained tasks. ([ACL Anthology][3])
The trade is exactly the one McNamara states: specialization for generality. You have not built a generally smarter 0.8B model. You have built something closer to an extremely practiced specialist. Ask it the task it has seen millions of variations of and it can be exceptional; move sufficiently outside that distribution and the giant generalist should recover its advantage.
The broader implication is stronger than “fine-tuning works.” For stable, repeated AI workloads, frontier intelligence can become a training-time dependency rather than an inference-time dependency. You pay for expensive reasoning when discovering and repairing failures, then amortize that intelligence into a small model whose cheap forward pass reproduces the learned behavior millions of times. That is the economic idea behind the entire thread.
[1]: https://x.noodl3.net/tobi/status/2094808564355191249?utm_source=chatgpt.com "tobi lutke (@tobi): \"Training tiny models for special purpose use cases works so incredibly well if you have a great self improving recursive flywheel. Shopify ML team is on fire. finetuned 0.8b model beats GPT 5.6-sol xhigh in this very specialized task.\" | nitter" [2]: https://shopify.engineering/sidekicks-continual-learning-loop?utm_source=chatgpt.com "Sidekick's continual learning loop (2026) - Shopify" [3]: https://aclanthology.org/2025.emnlp-main.9/?utm_source=chatgpt.com "Comparing Specialised Small and General Large Language Models on Text Classification: 100 Labelled Samples to Achieve Break-Even Performance - ACL Anthology"