--- name: agfx-writing-bindless-shaders description: ALWAYS use when writing or modifying HLSL shaders for AGFX, or wiring the agfxShaderCompiler / agfxShaderModule / agfxRenderPipeline / agfxComputePipeline pipeline around them. Trigger for AGFX_PUSH_CONSTANTS, ResourceHandle, AGFXTexture2D/AGFXRWTexture2D/AGFXStructuredBuffer/AGFXByteAddressBuffer/AGFXSampler, ResourceDescriptorHeap/SamplerDescriptorHeap, agfxCompileShader, agfxShaderCompilerOptions, agfxShaderModuleCreate, register(b0)/register(b1), "bindless", "push constants", writing a new .hlsl file for AGFX, mesh/task/compute shader entry points (main_vs/main_ps/main_cs/main_ms/main_as). Do NOT trigger for render pass/attachment authoring in host code — use agfx-render-targets-and-passes. Do NOT trigger for resource-state barriers/fences — use agfx-synchronization. Do NOT trigger for swap chain/present — use agfx-presentation-and-swapchain. --- # AGFX Bindless HLSL Shaders ## Overview AGFX shaders are HLSL (SM 6.6), compiled with DXC to a per-platform target: DXIL on Windows (`agfx_shader_compiler_windows.cpp`); DXIL then translated to Metal IR via the Metal shader converter on macOS (`agfx_shader_compiler_mac.mm`, using `IRRootSignatureFlagSamplerHeapDirectlyIndexed | IRRootSignatureFlagCBVSRVUAVHeapDirectlyIndexed`); SPIR-V on Linux (`agfx_shader_compiler_linux.cpp`, DXC's `-spirv` target via a `dlopen`'d `libdxcompiler.so` — path settable through `agfxShaderCompilerOptions::dxCompilerPath`, defaulting to `data/dlls/libdxcompiler.so`). Every target is **fully bindless, direct-indexed heaps**. The compiler defines `AGFX_METAL` on macOS and `AGFX_VULKAN` on Linux automatically; `data/shaders/agfx.h` uses these to hide the per-backend differences (e.g. `[[vk::push_constant]]`), so shader authors never branch on them for ordinary resource access. There is no per-draw descriptor table, no `register(t0, space0)` binding model, and no `Bind*` API on the C side beyond push constants. Every resource a shader touches — textures, buffers, samplers — is accessed by a `ResourceHandle` (a plain `uint` index) pulled out of `ResourceDescriptorHeap`/`SamplerDescriptorHeap` and wrapped in one of the `AGFX*` helper classes declared in `data/shaders/agfx.h`. The host side hands shaders these handles two ways: almost always via push constants (`agfxRenderPassPushConstants`/`agfxComputePassPushConstants`, bound at `register(b0)`), or, for structured scene/per-draw constant data, by putting the handle to a constant buffer *inside* the push constants and loading it as an `AGFXStructuredBuffer` in the shader (see `sceneCB` pattern below) rather than a second root CBV. ## Ownership **Owns:** - The bindless resource-access pattern: `ResourceHandle`, `AGFXTexture1D/2D/2DArray/3D/Cube`, `AGFXRWTexture1D/2D/3D`, `AGFXStructuredBuffer`/`AGFXRWStructuredBuffer`, `AGFXByteAddressBuffer`/`AGFXRWByteAddressBuffer`, `AGFXSampler`/`AGFXComparisonSampler`, `AGFXRaytracingAccelerationStructure` - Push constants: `AGFX_PUSH_CONSTANTS(type, name)` at `register(b0)` (a `[[vk::push_constant]]` block on Vulkan — the macro hides it), and the optional per-draw ID (`AGFX_DECLARE_DRAW_ID()`/`AGFX_DRAW_ID()`: `register(b1)` on D3D12/Metal, the SPIR-V `DrawIndex` builtin on Vulkan), which is how a shader replayed from an indirect bundle recovers which draw it is — but note the value's meaning diverges on Vulkan, see `agfx-mdi` - Entry point / stage conventions (`main_vs`, `main_ps`, `main_cs`, `main_ms`, `main_as`) and matching `agfxShaderModuleType` - `agfxShaderCompilerOptions`/`agfxShaderCompilerResult` and `agfxCompileShader` — the HLSL → DXIL (Windows) / Metal IR (macOS) / SPIR-V (Linux) pipeline - Wiring compiled `agfxShaderModule`s into `agfxRenderPipelineCreateInfo`/`agfxComputePipelineCreateInfo` **Doesn't own:** - Render pass/attachment setup the pipeline is later bound and drawn within → `agfx-render-targets-and-passes` - Barriers needed before a shader can safely read/write a resource (state transitions, UAV hazard barriers) → `agfx-synchronization` - The `AGFXIndirectDraw*Bundle` append helpers declared in the same header (`Create`/`Draw`/`DrawIndexed`/`DrawMesh`/`Dispatch`) and the GPU-driven submission model around them → `agfx-mdi` - Swap chain / back buffer acquisition → `agfx-presentation-and-swapchain` ## References The bindless helper header is a single shared file at `data/shaders/agfx.h`, included by every shader in the repo as `#include "data/shaders/agfx.h"` (a repo-root-relative path, not relative to the including shader) — **always `#include` it first** in a new shader and read it before inventing a new resource-access pattern; it is the complete list of what's available. Real shader examples: `data/shaders/demo/gbuffer.hlsl` (vertex+fragment, structured vertex pulling, textures+sampler), `data/shaders/demo/gbuffer_indirect.hlsl` (the GPU-driven variant, split into its own file for the Vulkan push-constant rule below), `data/shaders/demo/ssao.hlsl` (compute, RW texture output, scene CB), `data/shaders/demo/mipgen.hlsl` (minimal compute), `data/shaders/demo/deferred_lighting.hlsl`, `data/shaders/demo/shadow_depth.hlsl`, `data/shaders/demo/tonemap.hlsl`, `data/shaders/demo/imgui.hlsl`. Host-side compile+load pattern: `agfx_demo/deferred_renderer.cpp`'s `CompileShader` helper, `agfx_demo/ssao.cpp`, `agfx_demo/agfx_mipgen.cpp`. Compiler internals: `agfx_shader/agfx_shader_compiler.h` and the per-platform `agfx_shader_compiler_windows.cpp`/`_mac.mm`/`_linux.cpp`. ## Design Patterns ### Minimal shader skeleton ```hlsl #include "data/shaders/agfx.h" // one shared header for all shaders; path is repo-root-relative struct MyPushConstants { ResourceHandle someTex; ResourceHandle someSampler; ResourceHandle sceneCB; float someScalar; }; AGFX_PUSH_CONSTANTS(MyPushConstants, g_Constants); ``` `AGFX_PUSH_CONSTANTS` expands to `ConstantBuffer name : register(b0)` — this is the *only* resource binding declaration a shader normally needs. Everything else is created inline from a `ResourceHandle` field on that struct. ### Reading resources: create-from-handle, then use ```hlsl AGFXTexture2D albedoTex = AGFXTexture2D::Create(g_Constants.albedoTex); AGFXSampler samp = AGFXSampler::Create(g_Constants.textureSampler); float4 c = albedoTex.Sample(samp, uv); AGFXRWTexture2D aoOut = AGFXRWTexture2D::Create(g_Constants.aoTex); aoOut.Store(int2(id.xy), float4(ao, ao, ao, 1.0f)); AGFXStructuredBuffer vertices = AGFXStructuredBuffer::Create(g_Constants.vertexBuffer); SceneVertex v = vertices.Load(vertexID); ``` Pick the wrapper by both dimensionality and read/write need: `AGFXTexture2D` (read-only, sampled) vs `AGFXRWTexture2D` (read/write, `Load`/`Store` only, no filtering) — matching `AGFX_TEXTURE_USAGE_SAMPLED` vs `AGFX_TEXTURE_USAGE_STORAGE` on the host-side `agfxTextureCreateInfo`. Use `AGFXStructuredBuffer` for typed per-element buffer reads (scene constants, vertex pulling) and `AGFXByteAddressBuffer`/`AGFXRWByteAddressBuffer` for raw/untyped access — matching `AGFX_BUFFER_VIEW_TYPE_STRUCTURED` vs `AGFX_BUFFER_VIEW_TYPE_RAW` host-side. ### Scene/per-frame constants: no second CBV — nest a handle in push constants There is no root-level CBV beyond `b0`'s push constants. To pass a larger, per-frame constant buffer, put its `ResourceHandle` as a field in the push-constant struct and load it as a one-element `AGFXStructuredBuffer` inside the shader: ```hlsl struct GBufferSceneConstants { float4x4 viewProj; }; struct GBufferPushConstants { float4x4 worldMatrix; ResourceHandle sceneCB; // ... }; AGFX_PUSH_CONSTANTS(GBufferPushConstants, g_Constants); vs_out main_vs(uint vertexID : SV_VertexID) { AGFXStructuredBuffer sceneCB = AGFXStructuredBuffer::Create(g_Constants.sceneCB); GBufferSceneConstants scene = sceneCB.Load(0); // ... } ``` Host side, this `sceneCB` handle comes from `agfxBufferViewGetHandle` on an `agfxBufferView` created with `AGFX_BUFFER_VIEW_TYPE_CONSTANT` (or `STRUCTURED`, since the shader reads it as a structured buffer either way) over an upload-heap `agfxBuffer`. ### Vertex pulling instead of input-assembler vertex buffers AGFX vertex shaders don't use IA vertex attributes — they take `SV_VertexID` and manually pull from a structured buffer, since bindless makes an explicit vertex-buffer handle no more expensive than an IA binding and lets one pipeline draw meshes with arbitrary vertex layouts: ```hlsl vs_out main_vs(uint vertexID : SV_VertexID) { AGFXStructuredBuffer vertices = AGFXStructuredBuffer::Create(g_Constants.vertexBuffer); SceneVertex v = vertices.Load(vertexID + g_Constants.vertexOffset); // ... } ``` `vertexOffset` in push constants lets one shared vertex buffer serve multiple meshes/draws without rebinding. ### Compute shaders: bounds check, then dispatch-sized work ```hlsl struct MyPushConstants { ResourceHandle srcTex; ResourceHandle dstTex; uint dstWidth; uint dstHeight; }; AGFX_PUSH_CONSTANTS(MyPushConstants, g_Constants); [numthreads(8, 8, 1)] void main_cs(uint3 id : SV_DispatchThreadID) { if (id.x >= g_Constants.dstWidth || id.y >= g_Constants.dstHeight) return; // ... } ``` The `[numthreads(x, y, z)]` values must match the `groupSizeX/Y/Z` passed to `agfxComputePipelineCreateInfo` host-side (or, for mesh/task shaders, the reflected `meshSizeX/Y/Z`/`taskSizeX/Y/Z` the compiler extracts automatically — see below). Always bounds-check against actual target dimensions since dispatch group counts are typically rounded up. ### Entry point / stage naming convention Existing shaders use `main_vs`, `main_ps`, `main_cs`, and (for mesh pipelines) `main_ms`/`main_as`. Match this convention for new shaders and pass the matching `agfxShaderStage`/`agfxShaderModuleType` pair host-side — the DXC target profile (`vs_6_6`, `ps_6_6`, `cs_6_6`, `ms_6_6`, `as_6_6`) is derived from `agfxShaderStage` in each platform's `agfx_shader_compiler_*` file, so stage and entry point must agree. ### Host-side: compile → shader module → pipeline ```cpp // Typical helper (see deferred_renderer.cpp's CompileShader) agfxShaderCompilerOptions options = {}; options.stage = AGFX_SHADER_STAGE_FRAGMENT; strncpy(options.entryPoint, "main_ps", sizeof(options.entryPoint) - 1); options.sourceCode = source.data(); options.sourceCodeSize = (uint32_t)source.size(); agfxShaderCompilerResult result = {}; agfxCompileShader(&options, &result); agfxShaderModuleCreateInfo moduleInfo = {}; moduleInfo.code = result.compiledCode; moduleInfo.codeSize = result.compiledSize; moduleInfo.entryPoint = "main_ps"; moduleInfo.type = AGFX_SHADER_MODULE_TYPE_FRAGMENT; agfxShaderModule* fragmentModule = agfxShaderModuleCreate(device, &moduleInfo); ``` Compile vertex+fragment (or mesh[+task]) modules separately and attach both to `agfxRenderPipelineCreateInfo::vertexShader`/`fragmentShader` (or `meshShader`/`taskShader`); a single `computeShader` module goes into `agfxComputePipelineCreateInfo`. `agfxShaderModuleDestroy` is safe immediately after the pipeline(s) built from it are created — the module isn't referenced afterward. For mesh/task shaders, read `result.meshSizeX/Y/Z`/`taskSizeX/Y/Z` (populated by compiler reflection — Metal shader converter on macOS, SPIRV-Reflect on Linux — not something you hand-specify) and feed them into `agfxRenderPipelineCreateInfo::meshGroupSizeX/Y/Z`/`taskGroupSizeX/Y/Z` — these must match what the shader actually declares or dispatch counts silently disagree between backends. ### Common mistakes - Declaring a second `register(bN)`/`register(tN)`/`register(sN)` resource binding instead of routing everything through push constants + `ResourceDescriptorHeap`/`SamplerDescriptorHeap` — AGFX's binding surface is exactly the push constants (`b0`), the draw ID (`b1` on D3D12/Metal, a builtin on Vulkan), and direct-indexed heaps; anything else won't be bound. - **One `AGFX_PUSH_CONSTANTS` block per translation unit.** The SPIR-V backend rejects a file that declares more than one `[[vk::push_constant]]` block, even when only one entry point is selected for compilation — a file with multiple entry points that need *different* push-constant structs must be split into separate `.hlsl` files. This is why the demo's GPU-driven G-buffer path lives in `gbuffer_indirect.hlsl`, not in `gbuffer.hlsl`. - Calling `AGFX_DRAW_ID()` in a pixel shader. On Vulkan it resolves to the SPIR-V `DrawIndex` builtin, valid only in vertex/mesh/task stages — read it in the vertex shader and forward it through a `nointerpolation` interpolant (see `gbuffer_indirect.hlsl`). - Using `AGFXTexture2D` (read-only) where the texture was created with `AGFX_TEXTURE_USAGE_STORAGE` and needs `Store`, or vice versa — pick the wrapper matching the host-side `agfxTextureUsage`/`agfxTextureViewCreateInfo::writeable`. - Forgetting the bounds check in a compute shader before writing to an `AGFXRWTexture2D` sized smaller than `numthreads`-rounded dispatch dimensions. - Mismatching entry-point name and `agfxShaderStage`/`agfxShaderModuleType` between the compile options and the module create info — the DXC profile is derived from the stage, and a mismatch will misdirect the compiler or produce a module that doesn't bind to the intended pipeline slot.