Sutton observed in 2019 that general methods beat domain-specific ones when given enough compute and data. Chess engines that learned to play beat engines built on grandmaster knowledge. Speech recognizers trained on raw audio beat ones built on hand-tuned phoneme models.
In college I worked on analog compute circuits and resonator-based architectures. I still buy esoteric hardware platforms to see if I can eke out performance and efficiency that beats my trusty Apple with NPU and a gaming machine with NVIDIA GPU. My desk has enough of them that my wife's main challenge is finding space for her coffee cup without shorting something out. Each impulsive purchase eventually abandoned as the standard solutions won again. Which leads me to:
The hardware version of the Bitter Lesson:
Novel hardware platforms that require a custom computational primitive to realize their advantages get displaced by the dominant general purpose ones.
Three forces drive this. The GPU's own history shows all three.
Novel platforms compete against purpose-built and highly optimized inference hardware.
Benchmarks pit them against unoptimized GPU baselines, but real deployed workloads run on purpose-built and highly optimized inference hardware tuned at every level: silicon, memory, compiler, model. The dominant platform has decades of process node optimization, memory bandwidth engineering, and packaging investment behind it.
GPU optimization is funded by billions in datacenter revenue. Better compilers, kernels, and silicon are justified by real profit at scale. Novel hardware adopts these techniques defensively, spending scarce engineering resources chasing a moving target with no comparable revenue base.
Data, compression, features and embeddings co-evolve.
With most traditional signal domains already owned by the dominant platform, sparse signals are where novel hardware hopes to find a foothold. But those domains have the least data. The signals with the most data (audio, video, text) accumulated enough volume to justify decades of investment: first compression standards (mel spectrograms, video codecs, tokenization), then hand-engineered features, then learned representations, and now dense embeddings from foundation models. Each layer built on the last, deepening the dominant platform's hold. The sparsity advantage novel hardware was hoping to exploit was captured upstream, long before it arrived.
The domains without those compression standards are the domains without training data.
And for the hardware that does target the unexplored sparse regime, the training problem may be specific to those domains, rarely encountered before because the dominant platform never needed to train there.
The published results document what might be called an uncertainty principle for novel hardware: the sparsity levels that would make it genuinely efficient are the same levels at which training breaks down.
Sparse activations suppress the gradients needed to learn, and the hyperparameters that maximize accuracy are not the same ones that minimize hardware energy. I have not found a published result that resolves the tension on a non-trivial task.
Scale works against novel hardware.
The most efficient small models today are built by distilling large ones or training on synthetic data generated by them. Phi-3-mini reaches GPT-3.5 performance at 3.8B parameters partly because it trained on data generated by GPT-4. The same approach is available to novel hardware: generate synthetic training data from conventional large models. But in the sparse signal domains where novel hardware has its efficiency advantage, no foundational model exists to generate from. GPT-4 cannot synthesize DVS event streams at the scale needed to train competitive models. And as the dominant platform's large model collection grows, the gap between what can be distilled from it and what novel hardware can produce widens further.
The gap widens every year. This force didn't exist a decade ago.
The GPU and TPU follow the law, not an exception.
The GPU and TPU look like counterexamples but both follow the law. CPU came first and dominated compute, already running vector SIMD extensions (MMX in 1996, SSE in 1999) for graphics and multimedia workloads. GPU scaled the same SIMD concept to thousands of parallel lanes optimized for floating point matrix operations, on the already existing CMOS process. No new physics, no new computational primitive. Just more of the same vector math that gaming demanded at scale, which turned out to be exactly what deep learning needed. TPU took a different internal organization (systolic arrays instead of SIMD lanes) but the fundamental operation stayed the same: multiply-accumulate on matrices. CPU, GPU, TPU: three architectures, one primitive. That scale then produced the large models that became the foundation for distillation.
So where does that leave us?
These novel architectures rarely win the general market. The investment lead means they compete against decades of optimization they cannot match. The best domains for their efficiency claims turn out to be exactly the ones with no training data. And scale working against novel hardware means the gap widens over time.
I found no novel hardware architecture has escaped all three forces yet. But when one does, it will be on my desk promptly.
