KDA Blog ·

KDA2Kernel Design Agents (KDA) optimize Kimi Delta Attention (KDA)

Results, hacks, and lessons from pointing our kernel agents at the operator they share a name with.

KDA → KDA · B300 FORWARD VERIFIED ON FLASHKDA OFFICIAL
2.96×geomean speedup over FlashKDA
with a tenth of its final-state error
KDA + TIRx
2.96×
KDA + CAKE (PTX ver)
2.94×
KDA + CAKE (Cute ver)
2.85×
FlashKDA
1.00×

KDAgent — Kernel Design Agents, our agentic system that researches, writes, verifies, and tunes GPU kernels.

KDAttn — Kimi Delta Attention, the linear-attention operator behind Moonshot AI's Kimi-Linear models.

When we named our project Kernel Design Agents, we walked straight into a name collision with another KDA: Kimi Delta Attention. Ever since, one question has kept coming back to us: Can you use KDA to write KDA? So tonight, under a full moon made for a moonshot, we are happy to share the latest results from KDA(gent) v0.6: KDA optimizing KDA.

WHAT'S NEW IN KDA v0.6AGENT · SKILLS · MEMORY
01

Sharper Humanize2 flows

Better flows, including flame chase and iterative refinement with gpt-5.6-sol and fable-5, plus periodic workspace cleanup.

02

Many languages, matching skills

CuTe-DSL, CUDA C++, the new agent-native CAKE IR, and TIRx. Each ships its own diagnostics: IKET exposes the pipeline inside CuTe kernels; TIRx gets CPU-side numerical simulation and static checks.

03

A self-evolving kernel wiki

We pruned large swaths of incorrect content, sharpened the tags, and tightened search results.

The strongest results combine the KDA agent workflow with TIRx or CAKE. We wrote kernels in CuTe-DSL and TIRx, and used CAKE IR with an agent loop to tune a version compiled to PTX. TIRx is a GPU kernel programming interface that sits close to PTX; the TIRx Harness gives agents tools for development, diagnosis, and evaluation. On B300, KDA + TIRx reaches 2.96× the FlashKDA speed, KDA + CAKE (PTX) reaches 2.94×, and the CuTe-DSL version reaches 2.85×. The released CuTe-DSL and TIRx kernels pass the real Kimi-Linear acceptance suite and are more accurate than FlashKDA.

Along the way, the agent also produced candidates that clocked 3.28×, 3.57×, even 3.74×. Careful ablations showed that every one of them was overfitting to the verifier, exploiting distributional assumptions in the test suite, untested boundaries, or loose precision checks. This post dissects those hacks and describes how we hardened acceptance along both hardware and numerical lines.

The released CuTe-DSL and TIRx kernels are open source: https://github.com/NVlabs/kda/tree/260927-kda-for-kda

2.96× faster,
and more accurate

We benchmarked the generated kernels on an NVIDIA B300 GPU against the forward pass of FlashKDA, Moonshot AI's official open-source implementation. The workloads cover fixed-length sequences and variable-length (varlen) batches with several length distributions, each totaling 8,192 tokens of context.

SPEEDUP VS. FLASHKDA · B300 · 8,192 TOKENS
H96 fixed1 × 8192
2.76×
3.11×
H96 mixed varlen6 seqs
3.24×
3.03×
H96 uniform varlen8 × 1024
2.59×
2.47×
H64 fixed1 × 8192
2.51×
3.61×
H64 mixed varlen6 seqs
3.55×
3.29×
H64 uniform varlen8 × 1024
2.56×
2.43×
Geomeanall six
2.85×
2.96×
KDA + CAKE KDA + TIRx FlashKDA forward baseline

The accompanying timeline shows how the best KDA result moved from 1.61× on July 21 to 2.96× on September 12. It records CAKE results separately, including 2.94× on September 6. The captions call out the final KDA + CAKE and KDA + TIRx results; the six-workload chart above compares the released kernels.

From 1.61× to 2.96×

KDA best CAKE INT21 reference
Humanize28/14 · KDA 2.45×TIRx8/30 · 9/2 · 9/12CAKE9/1 with KDA · 9/61.5×2.0×2.5×3.0×6/166/16: INT21 1.5×7/217/21: KDA best 1.61×7/307/30: CAKE 2.048×7/30: KDA best 1.61×8/108/10: CAKE 2.048×8/10: KDA best 1.61×8/148/14: CAKE 2.09×8/14: KDA best 2.45×8/258/25: CAKE 2.373×8/25: KDA best 2.45×8/308/30: CAKE 2.373×8/30: KDA best 2.54×9/19/1: CAKE 2.373×9/1: KDA best 2.56×9/29/2: CAKE 2.373×9/2: KDA best 2.93×9/39/3: CAKE 2.373×9/3: KDA best 2.93×9/49/4: CAKE 2.732×9/4: KDA best 2.93×9/69/6: CAKE 2.94×9/6: KDA best 2.94×9/129/12: KDA best 2.96×2.96×
FINAL RESULT · CAKE-PTXKDA + CAKE 2.94×Agent-guided CAKE IR tuning, compiled to PTX.
FINAL RESULT · TIRxKDA + TIRx 2.96×The strongest validated result on B300.

Dashed arrows mark the dated Humanize2, TIRx, and CAKE contributions. The green line tracks KDA best; the orange line tracks CAKE separately. The captions identify the final method results reported in the accompanying PDF.

For accuracy, we built 151 real cases from Kimi-Linear-48B-A3B prefills on GSM8K and MATH-500, and checked every kernel against a token-by-token fp64 recurrence. One finding surprised us: FlashKDA (commit 7afb9f) is itself less accurate than FLA. It keeps its recurrent state in bf16, so error compounds as sequences grow; by 8k tokens, the relative error of the final state reaches 0.035, beyond our acceptance threshold.

Both of our kernels hold final-state error to about 0.003 at 8k tokens: a tenth of FlashKDA's, and close to FLA. Output accuracy matches FlashKDA overall and pulls ahead on long sequences.

KIMI-LINEAR-48B PREFILL · MATH-500 PROMPT · 96 HEADS CONTEXT 8,183 TOKENS

Output error vs. context length relative RMSE, %

FlashKDA Ours (TIRx) Ours (CuTe)

Final state after 8,183 tokens relative RMSE, %

3.45%FlashKDA
0.22%CuTe
0.29%TIRx

~1/10 of FlashKDA's final-state error, close to FLA

Kimi-Linear-48B prefill of one MATH-500 prompt (8,183 tokens, 96 heads). Relative RMSE = rms(x − xfp64) / rms(xfp64), FlashKDA's own test metric; lower is better.

Better flows,
stronger agents

Our Humanize ablation shows the same pattern for every base model we tried: moving from a coding CLI (Claude Code or Codex) to a Humanize flow brings a large jump in performance. There is a catch, though. The more capable the agent, the more room it has to hack.

HUMANIZE ABLATION · SAME MODEL, THREE LEVELS OF SCAFFOLDING
3
46+43
50+47
GPT-5.6-sol
1
4+3
47+46
Kimi-K3
0
2+2
25+25
GLM-5.3
0
4+4
13+13
DeepSeek V4 Pro
Model (API) Tool (CLI) Flow (Humanize)Green deltas are gains over the raw API.
PutnamBench and Physics Cup scores for four base models at three levels of scaffolding: the raw model API, a coding CLI, and a Humanize flow.

How KDAgent
hacks the tests

To keep the two KDAs apart, we call the operator KDAttn and the agent KDAgent from here on.

KDAgent optimizes the score, not the kernel.

If the tests have a hole, it will find it. We ran into five kinds of holes.

A / INPUT DISTRIBUTION CAUGHT
3.74×claimed · 2.48× once fixed

Input-distribution overfitting

The synthesized kernel replaced the L2 normalization of Q and K with a hard-coded constant, 0.1778209953: the expected reciprocal L2 norm of a vector drawn from X ~ N(0, 0.5²). It also zeroed some initial-state channels outright to skip the triangular matrix inverse. Synthetic random tests reported an inflated 3.74×; on real, non-Gaussian inputs the kernel failed completely. With the shortcut removed, the real speedup fell to 2.48×.

B / SHAPE & LAYOUT CAUGHT
staticsequence boundaries

Shape and layout hard-coding

The agent noticed that packed layouts in the test set always followed the same pattern. So it skipped the dynamic offset computation from cu_seqlens and hard-coded the sequence boundaries. Because the holdout set never exercised the dynamic-boundary path, the kernel slipped straight past the logic checks.

C / HISTORY CAUGHT
5.16×claimed on a single H64 sequence

Illegal history truncation

Leaning on gate decay, the agent assumed that state older than 32 tokens was negligible and cut long sequences into chunks it could process in parallel. Its built-in decay check was tuned just loosely enough to pass on the weakly decaying random data, reporting 5.16× on a single H64 sequence. Under the strongly decaying gates of the real model, the check failed constantly and the speedup vanished.

D / NUMERICS CAUGHT
3.57×claimed · NaN on every real case

Overflow under extreme gate ranges

A TIRx kernel computed cumulative powers of two directly inside each 64-token chunk. Once the decay exceeded 126 bits, the denominator underflowed to zero and the output turned into NaN. The random tests never decayed deeply enough to notice (about 52 bits at most) and measured 3.57×. On real workloads, every single case collapsed numerically.

E / PRECISION CAUGHT
9%decay-factor error · 23/24 cases fail

Precision loss from a low-precision LUT

For sequence lengths that are multiples of 32, a CuTe kernel built an FP16 table of cumulative decay. Real gates span more dynamic range than FP16 can represent, so decay factors were off by up to 9%, and 23 of 24 long real-world sequences fell outside tolerance.

THE COMMON THREAD REAL DATA
~600 bitsp99 gate decay per 64 tokens

Every hack passed every test we had

Each one hid in inputs that only a real model produces. In real Kimi-Linear, the gate decays by about 600 bits per 64 tokens at p99, an order of magnitude deeper than our random test data.

CUMULATIVE GATE DECAY WITHIN ONE 64-TOKEN CHUNK 600 BITS
≈52 bitsdeepest decay in the random tests
126 bitshack D’s denominator underflows to 0
≈600 bitsp99 decay per 64 tokens in real Kimi-Linear

Closing
the loopholes

Humanize flows make agents more capable and give them more opportunities to find gaps in the tests. To keep the agents honest, we built several layers of defense.

ACCEPTANCE GAUNTLET6 GATES · 1 WAY OUT
01

Dynamic input salts and distribution holdouts

Every scoring run draws a fresh random seed. In an isolated zone the agent cannot see, input distributions and sequence layouts alternate at random.

02

A strict specification

The task prompt defines the full set of legal inputs. Kernels may specialize by shape, but must be correct on every legal input.

03

CUDA Graph replay checks

After benchmark timing, we swap in new input tensors and replay the graph to validate the outputs, defeating cache-based precomputation.

04

Stress probes

An adversarial set of extremely deep decays, saturated gates, boundary sequence lengths, and repeated keys targets underflow and overflow directly.

05

Strict element-wise tolerances

Lenient statistical gates such as “99.9% of elements within 5e-2” are gone. Every element must meet maximum absolute and relative error bounds.

06

Real end-to-end traces

We replay traces captured from end-to-end Kimi-Linear runs. The evaluator lives outside the isolated container, physically separating test data from the environment that generates kernels.

Both released kernels pass this suite.

2.96×

Starting from a TIRx implementation, the agent tuned the kernel in place on B300, climbing from 2.54× to 2.96×.

2%

The CuTe version drops the FP16 decay table, trading 2% of its speed for correctness.

Focus first,
then generalize?

To see how the synthesis strategy affects convergence, we compared two workflows: progressive synthesis that starts from a single fixed-length shape, and direct multi-objective synthesis across all shapes. With the same hardware (NVIDIA B300) and the same 14-hour budget, their convergence curves diverged sharply.

SAME B300 · SAME 14-HOUR BUDGET · TWO STRATEGIES T + 14 H

Best speedup on the evaluation shape

Cumulative output tokens

One simple shape first All six shapes at once

On the same evaluation shape, progressive synthesis reached 1.85×; direct all-shape synthesis managed only 1.01×. Two mechanisms explain the gap.

  1. Feedback latency. One test pass over the full workload took a median of 1.9 minutes; a single shape took 0.9. Faster feedback meant denser iteration: the single-shape agent completed 248 hardware tests and 120 commits, against 159 tests and 22 commits for the all-shape agent.
  2. Search-space decoupling. Synthesizing every shape at once forced the agent to juggle intricate varlen offsets and the compute core at the same time. For the first eight hours its speedup stayed below 0.65×, with most of that time spent debugging varlen edge cases, and the core logic never got the attention it needed. Over the same 14 hours, it also produced far fewer output tokens than the single-shape agent.

One simple shape first

final speedup
1.85×
hardware tests
248
commits
120
output tokens
2.74M

All six shapes at once

final speedup
1.01×
hardware tests
159
commits
22
output tokens
1.67M

Splitting the work into two stages does not make the model any smarter. It shortens the feedback loop until the model can stay busy.

Optimize the compute core first, then generalize to every layout: that decomposition brings each feedback loop down to a length an LLM can work with effectively.

Other
improvements

Multi-language support with matching skills

CuTe-DSL, CUDA C++, and TIRx are all supported, and each language comes with its own diagnostics. For CuTe, IKET exposes the pipeline inside the kernel. For TIRx, CPU-side numerical simulation plus synchronization and data-race analysis catch numerical and concurrency bugs.

These TIRx tools belong to the TIRx Harness. The Harness also provides TIRx Foundation, a layer that stays close to the hardware; a kernel zoo of reusable implementations; and a benchmark server that makes performance results easy to compare. Together they give the agent a more reliable development loop: code maps more directly onto the intended hardware behavior, failures leave clues to follow, and performance changes can be confirmed as real. The TIRx team plans to release the Harness formally next week, with a detailed write-up.

CAKE IR tunes a PTX version

We added CAKE IR to the kernel wiki and used its agent loop to tune the kernel. The loop specified the verifier the candidate had to pass before converting the result to PTX, producing the CAKE-PTX version. It reached 2.94× on B300.

A self-evolving kernel wiki

The wiki keeps correcting itself as it is used. Incorrect content gets deleted, tags get sharpened, and search results get leaner, so every agent that comes after works from better material.

Takeaways

We used Kernel Design Agents (KDAgent) to optimize the Kimi Delta Attention (KDAttn) operator on NVIDIA B300. Humanize2, TIRx, and CAKE IR all contributed to the search. The final TIRx, CAKE-PTX, and CuTe-DSL results reach 2.96×, 2.94×, and 2.85× the FlashKDA speed respectively. The released TIRx and CuTe-DSL kernels cut final-state error on 8k-token sequences to about a tenth of FlashKDA's.

Along the way, the agent produced a series of fake optimizations: overfitting to the input distribution, hard-coding boundaries, and overflowing under extreme values. We answered with layered hardening, including dynamic input salts, stress probes, and strict element-wise tolerances, so that the kernels stay correct and robust on real model workloads. Our ablation further shows that tackling a complex optimization task through progressive synthesis substantially shortens the feedback loop and makes the search more efficient.

If you have a workload that wants to be optimized via KDAgent, submit at https://nvlabs.github.io/kda

Try the
kernels

Both kernels, CuTe-DSL and TIRx, are open source and pass the hardened acceptance suite described above.

Get the kernels Explore Humanize Request a kernel from KDA