KDA2Kernel Design Agents (KDA) optimize Kimi Delta Attention (KDA)
Results, hacks, and lessons from pointing our kernel agents at the operator they share a name with.
with a tenth of its final-state error
KDAgent — Kernel Design Agents, our agentic system that researches, writes, verifies, and tunes GPU kernels.
KDAttn — Kimi Delta Attention, the linear-attention operator behind Moonshot AI's Kimi-Linear models.
When we named our project Kernel Design Agents, we walked straight into a name collision with another KDA: Kimi Delta Attention. Ever since, one question has kept coming back to us: Can you use KDA to write KDA?
So tonight, under a full moon made for a moonshot, we are happy to share the latest results from KDA(gent) v0.6: KDA optimizing KDA.
Sharper Humanize2 flows
Better flows, including flame chase and iterative refinement with gpt-5.6-sol and fable-5, plus periodic workspace cleanup.
Many languages, matching skills
CuTe-DSL, CUDA C++, the new agent-native CAKE IR, and TIRx. Each ships its own diagnostics: IKET exposes the pipeline inside CuTe kernels; TIRx gets CPU-side numerical simulation and static checks.
A self-evolving kernel wiki
We pruned large swaths of incorrect content, sharpened the tags, and tightened search results.
The strongest results combine the KDA agent workflow with TIRx or CAKE. We wrote kernels in CuTe-DSL and TIRx, and used CAKE IR with an agent loop to tune a version compiled to PTX. TIRx is a GPU kernel programming interface that sits close to PTX; the TIRx Harness gives agents tools for development, diagnosis, and evaluation. On B300, KDA + TIRx reaches 2.96× the FlashKDA speed, KDA + CAKE (PTX) reaches 2.94×, and the CuTe-DSL version reaches 2.85×. The released CuTe-DSL and TIRx kernels pass the real Kimi-Linear acceptance suite and are more accurate than FlashKDA.
Along the way, the agent also produced candidates that clocked 3.28×, 3.57×, even 3.74×. Careful ablations showed that every one of them was overfitting to the verifier, exploiting distributional assumptions in the test suite, untested boundaries, or loose precision checks. This post dissects those hacks and describes how we hardened acceptance along both hardware and numerical lines.
The released CuTe-DSL and TIRx kernels are open source: https://github.com/NVlabs/kda/tree/260927-kda-for-kda
01 · RESULTS
2.96× faster,
and more accurate
We benchmarked the generated kernels on an NVIDIA B300 GPU against the forward pass of FlashKDA, Moonshot AI's official open-source implementation. The workloads cover fixed-length sequences and variable-length (varlen) batches with several length distributions, each totaling 8,192 tokens of context.
The accompanying timeline shows how the best KDA result moved from 1.61× on July 21 to 2.96× on September 12. It records CAKE results separately, including 2.94× on September 6. The captions call out the final KDA + CAKE and KDA + TIRx results; the six-workload chart above compares the released kernels.
From 1.61× to 2.96×
Dashed arrows mark the dated Humanize2, TIRx, and CAKE contributions. The green line tracks KDA best; the orange line tracks CAKE separately. The captions identify the final method results reported in the accompanying PDF.
For accuracy, we built 151 real cases from Kimi-Linear-48B-A3B prefills on GSM8K and MATH-500, and checked every kernel against a token-by-token fp64 recurrence. One finding surprised us: FlashKDA (commit 7afb9f) is itself less accurate than FLA. It keeps its recurrent state in bf16, so error compounds as sequences grow; by 8k tokens, the relative error of the final state reaches 0.035, beyond our acceptance threshold.
Both of our kernels hold final-state error to about 0.003 at 8k tokens: a tenth of FlashKDA's, and close to FLA. Output accuracy matches FlashKDA overall and pulls ahead on long sequences.
Output error vs. context length relative RMSE, %
Final state after 8,183 tokens relative RMSE, %
~1/10 of FlashKDA's final-state error, close to FLA
02 · HUMANIZE
Better flows,
stronger agents
Our Humanize ablation shows the same pattern for every base model we tried: moving from a coding CLI (Claude Code or Codex) to a Humanize flow brings a large jump in performance. There is a catch, though. The more capable the agent, the more room it has to hack.
03 · REWARD HACKING
How KDAgent
hacks the tests
To keep the two KDAs apart, we call the operator KDAttn and the agent KDAgent from here on.
KDAgent optimizes the score, not the kernel.
If the tests have a hole, it will find it. We ran into five kinds of holes.
Input-distribution overfitting
The synthesized kernel replaced the L2 normalization of Q and K with a hard-coded constant, 0.1778209953: the expected reciprocal L2 norm of a vector drawn from X ~ N(0, 0.5²). It also zeroed some initial-state channels outright to skip the triangular matrix inverse. Synthetic random tests reported an inflated 3.74×; on real, non-Gaussian inputs the kernel failed completely. With the shortcut removed, the real speedup fell to 2.48×.
Shape and layout hard-coding
The agent noticed that packed layouts in the test set always followed the same pattern. So it skipped the dynamic offset computation from cu_seqlens and hard-coded the sequence boundaries. Because the holdout set never exercised the dynamic-boundary path, the kernel slipped straight past the logic checks.
Illegal history truncation
Leaning on gate decay, the agent assumed that state older than 32 tokens was negligible and cut long sequences into chunks it could process in parallel. Its built-in decay check was tuned just loosely enough to pass on the weakly decaying random data, reporting 5.16× on a single H64 sequence. Under the strongly decaying gates of the real model, the check failed constantly and the speedup vanished.
Overflow under extreme gate ranges
A TIRx kernel computed cumulative powers of two directly inside each 64-token chunk. Once the decay exceeded 126 bits, the denominator underflowed to zero and the output turned into NaN. The random tests never decayed deeply enough to notice (about 52 bits at most) and measured 3.57×. On real workloads, every single case collapsed numerically.
Precision loss from a low-precision LUT
For sequence lengths that are multiples of 32, a CuTe kernel built an FP16 table of cumulative decay. Real gates span more dynamic range than FP16 can represent, so decay factors were off by up to 9%, and 23 of 24 long real-world sequences fell outside tolerance.
Every hack passed every test we had
Each one hid in inputs that only a real model produces. In real Kimi-Linear, the gate decays by about 600 bits per 64 tokens at p99, an order of magnitude deeper than our random test data.
04 · HARDENING
Closing
the loopholes
Humanize flows make agents more capable and give them more opportunities to find gaps in the tests. To keep the agents honest, we built several layers of defense.
Dynamic input salts and distribution holdouts
Every scoring run draws a fresh random seed. In an isolated zone the agent cannot see, input distributions and sequence layouts alternate at random.
A strict specification
The task prompt defines the full set of legal inputs. Kernels may specialize by shape, but must be correct on every legal input.
CUDA Graph replay checks
After benchmark timing, we swap in new input tensors and replay the graph to validate the outputs, defeating cache-based precomputation.
Stress probes
An adversarial set of extremely deep decays, saturated gates, boundary sequence lengths, and repeated keys targets underflow and overflow directly.
Strict element-wise tolerances
Lenient statistical gates such as “99.9% of elements within 5e-2” are gone. Every element must meet maximum absolute and relative error bounds.
Real end-to-end traces
We replay traces captured from end-to-end Kimi-Linear runs. The evaluator lives outside the isolated container, physically separating test data from the environment that generates kernels.
Both released kernels pass this suite.
TIRX · TUNED IN PLACE ON B300
2.96×Starting from a TIRx implementation, the agent tuned the kernel in place on B300, climbing from 2.54× to 2.96×.
CUTE-DSL · FP16 TABLE REMOVED
2%The CuTe version drops the FP16 decay table, trading 2% of its speed for correctness.
05 · ABLATION
Focus first,
then generalize?
To see how the synthesis strategy affects convergence, we compared two workflows: progressive synthesis that starts from a single fixed-length shape, and direct multi-objective synthesis across all shapes. With the same hardware (NVIDIA B300) and the same 14-hour budget, their convergence curves diverged sharply.
Best speedup on the evaluation shape
Cumulative output tokens
On the same evaluation shape, progressive synthesis reached 1.85×; direct all-shape synthesis managed only 1.01×. Two mechanisms explain the gap.
- Feedback latency. One test pass over the full workload took a median of 1.9 minutes; a single shape took 0.9. Faster feedback meant denser iteration: the single-shape agent completed 248 hardware tests and 120 commits, against 159 tests and 22 commits for the all-shape agent.
- Search-space decoupling. Synthesizing every shape at once forced the agent to juggle intricate varlen offsets and the compute core at the same time. For the first eight hours its speedup stayed below 0.65×, with most of that time spent debugging varlen edge cases, and the core logic never got the attention it needed. Over the same 14 hours, it also produced far fewer output tokens than the single-shape agent.
One simple shape first
- final speedup
- 1.85×
- hardware tests
- 248
- commits
- 120
- output tokens
- 2.74M
All six shapes at once
- final speedup
- 1.01×
- hardware tests
- 159
- commits
- 22
- output tokens
- 1.67M
Splitting the work into two stages does not make the model any smarter. It shortens the feedback loop until the model can stay busy.
Optimize the compute core first, then generalize to every layout: that decomposition brings each feedback loop down to a length an LLM can work with effectively.
06 · ALSO IN v0.6
Other
improvements
Multi-language support with matching skills
CuTe-DSL, CUDA C++, and TIRx are all supported, and each language comes with its own diagnostics. For CuTe, IKET exposes the pipeline inside the kernel. For TIRx, CPU-side numerical simulation plus synchronization and data-race analysis catch numerical and concurrency bugs.
These TIRx tools belong to the TIRx Harness. The Harness also provides TIRx Foundation, a layer that stays close to the hardware; a kernel zoo of reusable implementations; and a benchmark server that makes performance results easy to compare. Together they give the agent a more reliable development loop: code maps more directly onto the intended hardware behavior, failures leave clues to follow, and performance changes can be confirmed as real. The TIRx team plans to release the Harness formally next week, with a detailed write-up.
CAKE IR tunes a PTX version
We added CAKE IR to the kernel wiki and used its agent loop to tune the kernel. The loop specified the verifier the candidate had to pass before converting the result to PTX, producing the CAKE-PTX version. It reached 2.94× on B300.
A self-evolving kernel wiki
The wiki keeps correcting itself as it is used. Incorrect content gets deleted, tags get sharpened, and search results get leaner, so every agent that comes after works from better material.
07 · CONCLUSION
Takeaways
We used Kernel Design Agents (KDAgent) to optimize the Kimi Delta Attention (KDAttn) operator on NVIDIA B300. Humanize2, TIRx, and CAKE IR all contributed to the search. The final TIRx, CAKE-PTX, and CuTe-DSL results reach 2.96×, 2.94×, and 2.85× the FlashKDA speed respectively. The released TIRx and CuTe-DSL kernels cut final-state error on 8k-token sequences to about a tenth of FlashKDA's.
Along the way, the agent produced a series of fake optimizations: overfitting to the input distribution, hard-coding boundaries, and overflowing under extreme values. We answered with layered hardening, including dynamic input salts, stress probes, and strict element-wise tolerances, so that the kernels stay correct and robust on real model workloads. Our ablation further shows that tackling a complex optimization task through progressive synthesis substantially shortens the feedback loop and makes the search more efficient.
If you have a workload that wants to be optimized via KDAgent, submit at https://nvlabs.github.io/kda
RUN IT YOURSELF
Try the
kernels
Both kernels, CuTe-DSL and TIRx, are open source and pass the hardened acceptance suite described above.
Get the kernels Explore Humanize Request a kernel from KDA