@waterloo_intern
Post
two weeks ago i went on @swyx's pod and said some things that i... should not have said.
a lot has happened since then, i owe you all an apology.
i'm sorry that i was right about every single thing.
a) re megakernels are dead why are megakernels useful? you spend two months writing a kernel to save time on launch overhead and poor inter-kernel overlap. you had PDL but then people said it wasn't perfect, that you could still get some marginal gains due to straggler CTAs and therefore- wait, sorry, I forgot, Rubin fixes that (kernel two needs 10 CTAs and kernel one has seven finished and three straggling, kernel two launches seven of its CTAs). given a long enough timeline, it all evens out. no serious inference provider is using a 67k loc hand-fused forward pass kernel in production, and the teams doing that are doing so out of pure research. dead.
b) re ASICs are dead i'm sorry. to be specific: data-center transformer-inference ASIC companies (not naming any) who etched the arch into silicon have bet on architectural convergence. read kimi's architecture. read deepseek. qwen. we did not converge, and probably will not. dead.
c) re gpu kernel dev is dead this one kind of hurts because it is (was) my job. gpu kernel optimization is the single most RL-able task in existence correct=check\_correctness(kernel, shape) for shape in shapes if all(correct): time(kernel) give an agent ncu cli and an mcp with nvidia's tribal knowledge and it's done. dead.
d) re NVIDIA is scared of AMD humans hate programming AMD. i'm sorry. it's just true. fine taking a performance hit as long as i don't have to touch rocm or a programming paradigm that says a warp is 64 threads (wtf?)...but an agent does not... so assuming software no longer moat, HBM capacity and bandwidth matter, and currently on perf / price they're goated. 'bUt NvIdIa iS gOaTeD oN hArDwArE sOfTwArE cOdEsIgN' and that's the new moat. watch how much tooling they open source to get kernel devs on nvidia.
apologies all.

Image from X post

Image from X post

Image from X post
Quoted post by Latent.Space (@latentspacepod) The Inference Engineering Masterclass: 10x faster models, quantization, speculative decoding, Rubin, & self-optimizing AI https://www.latent.space/p/inference-eng
@Baseten @philipkiely and @waterloo\_intern explain what actually happens after a model is trained, why turning weights into a fast and reliable product creates an entirely new optimization problem, how quantization errors can cancel out to unlock more throughput, why inference teams are still finding 20–200% performance gains, how video generation runs into a quadratic attention wall, and how GLM-5.2 helped rewrite and optimize the GPU kernels serving GLM-5.2 itself.
![]()
Video thumbnail from X post
Explanation
What it says Ali argues four inference-engineering bets are becoming obsolete: megakernels, transformer-specific datacenter ASICs, hand-written GPU kernel optimization, and NVIDIA’s software moat over AMD. His thesis is that hardware/runtime improvements plus coding agents will erase much of the value of heroic manual optimization.
Context For megakernels, the key claim is about NVIDIA Rubin: finer-grained synchronization between dependent kernels can let already-ready CTAs from a consumer kernel start while straggler CTAs from the producer are still finishing. That reduces the launch/overlap penalty megakernels were partly designed to eliminate. For ASICs, he points to the divergence among Kimi, DeepSeek and Qwen architectures: hard-wiring “the transformer architecture” is dangerous if architectures keep changing. For kernel development, he thinks optimization is unusually amenable to RL/agents because correctness and runtime are mechanically measurable.
Why it matters The broader bet is that inference advantage shifts upward from bespoke kernels/silicon toward adaptable general-purpose GPUs, memory bandwidth/capacity, compiler/runtime machinery, and autonomous optimization agents. If true, AMD gains strategically because agents care less about ROCm ergonomics than humans do. The claims are intentionally categorical (“dead”); the material does not establish that production megakernels or ASICs literally disappear.
Images One diagram contrasts GPT-2 attention with increasingly exotic DeltaNet/Kimi architectures; another illustrates Rubin eliminating straggler-induced idle time through finer dependency tracking.