--- name: triton-kernel-writing description: Write or review Triton kernels for vLLM, with practical guidance for generated-code inspection, launch grids, indexing, specialization, tuning, and representative performance validation. --- # Triton Kernel Writing ## Implementation - Follow the official [Triton semantics](https://triton-lang.org/main/python-api/triton-semantics.html). Check it when behavior may differ from Python or NumPy, especially type promotion, integer division and modulo, casts, broadcasting, and variable scoping. - Use the Triton kernel generated by `torch.compile` as a possible implementation to inspect. Print Inductor's generated code with `TORCH_LOGS="output_code" .venv/bin/python