l8r Twats Library

@tobi

Post

Btw we open sourced the core infra piece that makes these self improving loops possible. https://tangleml.com

Quoted post by tobi lutke (@tobi) Training tiny models for special purpose use cases works so incredibly well if you have a great self improving recursive flywheel. Shopify ML team is on fire.

finetuned 0.8b model beats GPT 5.6-sol xhigh in this very specialized task.

Image from X post

Image from X post

Open quoted post on X

Explanation

Tobi Lütke, Shopify’s co-founder and CEO, is pointing at a very specific production result: Shopify took a tiny 0.8-billion-parameter Qwen3.5 model, fine-tuned it for one narrow job—generating “buyer profiles”—and got it to score slightly better on Shopify’s internal evaluator than GPT-5.6 Sol running at its highest reasoning setting: 84.6 vs. 83.0. At the same time, they cut the instruction prompt from 9.1K tokens to 1.1K and increased throughput from about 2 million profiles/day on their previous 2B model to 72 million/day on 100 H100s. Lütke’s follow-up is that Shopify has open-sourced Tangle, the ML pipeline infrastructure they use to make this kind of iterative training workflow manageable. ([Shopify][1])

The important qualification is that “beats GPT-5.6” does not mean the 0.8B model is generally smarter. It means it beats the general model on one highly constrained distribution with one internal scoring function. Think of it less like “a pocket calculator became smarter than a mathematician” and more like “we compiled one repetitive piece of the mathematician’s work into an extremely efficient specialized program.” The tiny model can devote essentially all of its representational capacity to the exact input→output mapping Shopify cares about.

The “self-improving recursive flywheel” is the central idea. The slide shows successive versions improving from 75.3 with 29K training examples, to 78.1 with 42K, to 84.6 with 54K. The likely loop is roughly: use a strong expensive model and/or human judgments to produce or score examples → collect the good/corrected examples → fine-tune the small model → evaluate it → deploy the better model → collect more production examples → repeat. Each generation of the model therefore helps produce data for the next generation. That is the recursion; “self-improving” is slightly loose language because humans, judges, data-selection logic, and larger models are still part of the system. ([Buttondown][2])

The 9.1K → 1.1K “gisted” prompt is almost as interesting as the model score. A generic model needs lots of instructions on every invocation: definitions, edge cases, formatting rules, examples, policy, and so on. Fine-tuning can move much of that behavior into the weights. You stop repeatedly transmitting thousands of prompt tokens and instead teach the model the task once through gradient updates. “Gisting” here means retaining a compact runtime representation/instruction set rather than spelling the full prompt out every time. That reduces both prefill computation and latency.

Why the 36× throughput increase can be much larger than the parameter reduction alone is that several gains multiply: 0.8B vs. 2B model size, a dramatically shorter prompt, presumably highly optimized batching, and a workload where the output itself is constrained. At tens of millions of runs per day, saving even a few thousand input tokens per request is enormous.

Tangle is not itself the clever learning algorithm. It is Shopify’s open-source orchestration layer for constructing and repeatedly running these ML/data pipelines: components, distributed execution, caching, artifact tracking, reproducibility, evaluation stages, retraining stages, etc. Shopify describes it as the “glue” connecting arbitrary code and ML steps, with content-based caching so unchanged stages need not be recomputed. ([Shopify][1])

The broader point Lütke is making is therefore quite strong: for stable, high-volume business tasks with a good evaluator and lots of examples, repeatedly calling a frontier model may be economically irrational. Use the frontier model as a teacher/data engine, distill the behavior into a tiny specialist, and keep recycling production experience into subsequent fine-tunes. The hard part is no longer merely “which base model is smartest?”; it is building the data-and-evaluation flywheel that lets the specialist keep improving.

[1]: https://shopify.engineering/tangle?utm_source=chatgpt.com "Tangle: An open-source ML experimentation platform built for scale (2025) - Shopify" [2]: https://buttondown.com/practical-ai/archive/webmcp-three-tools-that-make-a-rails-app-agent-ready/?utm_source=chatgpt.com "WebMCP: three tools that make a Rails app agent-ready • Buttondown"