l8r Twats Library

@antirez

Post

Follow my simple reasoning here. Large sparse LLMs are more powerful then dense one \while\ being faster and thus more energy efficient. Dense models are a need that is artificially created by VRAM scarcity. Today they are practically useful, but not for long.

Explanation

What it says antirez argues that large sparse LLMs—models that activate only part of their parameters per token, typically MoE-style—can deliver more total model capacity while being faster and more energy-efficient than similarly capable dense models. His stronger claim is that dense models persist mainly because GPU VRAM is scarce, not because dense architectures are intrinsically preferable.

Context This is a compressed architectural argument, not evidence by itself. The missing pieces are what “more powerful” means, what sparse architecture/model sizes he has in mind, and whether he is comparing equal parameter count, active parameter count, memory footprint, training cost, inference latency, or quality. Those choices can reverse the conclusion.

Why it matters The useful prediction is that as memory capacity/bandwidth constraints ease, frontier and local models should increasingly favor very large sparse models over smaller dense ones: huge stored parameter counts, but relatively few active per token. If correct, VRAM-heavy systems become valuable not because every parameter must compute simultaneously, but because they can store much larger expert pools. The claim is worth revisiting when comparing future dense vs. MoE models on quality per active FLOP, memory requirements, and real inference throughput.