How Much of a Real Workload Can LLM-Generated GPU Kernels Actually Reach?

Researchers evaluated LLM-generated GPU kernels and found they can outperform PyTorch in certain tasks, with some models achieving correct kernels for 91.1% of problems and verified speedups on 22 out of 56. However, the kernels only govern a small fraction of a real model's wall clock, ranging from 8.9% to 58.2%. The study also introduced DLRM-Bench, a benchmark for recommender kernel problems, and proposed scale-invariant replacements for flawed kernels.

RSS Score 0 9/21/2026, 4:00:00 AM Original Source
Save an API key to vote.