TypeGPU introduces a custom API to easily define and execute render and compute pipelines.
It abstracts away the standard WebGPU procedures to offer a convenient, type-safe way to run shaders on the GPU.
Creates a compute pipeline that executes the given callback in an exact number of threads.
This is different from createComputePipeline() in that it does a bounds check on the
thread id, where as regular pipelines do not and work in units of workgroups.
@param ― callback A function converted to WGSL and executed on the GPU.
It can accept up to 3 parameters (x, y, z) which correspond to the global invocation ID
of the executing thread.
@example
If no parameters are provided, the callback will be executed once, in a single thread.
const fooPipeline = root
.createGuardedComputePipeline(()=> {
'use gpu';
console.log('Hello, GPU!');
});
fooPipeline.dispatchThreads();
// [GPU] Hello, GPU!
@example
One parameter means n-threads will be executed in parallel.
The createRenderPipeline method creates a render pipeline by accepting an options object that specifies the vertex function, fragment function, targets, and optional additional settings.
vertex: The TgpuVertexFn or 'use gpu' callback to use as the vertex shader.
fragment: The TgpuFragmentFn or 'use gpu' callback to use as the fragment shader.
targets: A record defining the formats and behaviors of the color targets, similar to WebGPU’s GPUColorTargetState, but as a record with named targets.
depthStencil (optional): Depth-stencil state, same as WebGPU’s GPUDepthStencilState.
multisample (optional): Multisample state, same as WebGPU’s GPUMultisampleState.
primitive (optional): Primitive state, same as WebGPU’s GPUPrimitiveState.
The vertex function’s input parameters (non-builtin) are matched to vertex attributes specified in the pipeline’s vertex layout when executing. Vertex attributes are validated at the type level for compatibility.
Using the pipelines should ensure the compatibility of the vertex output and fragment input on the type level.
These parameters are identified by their names, not by their numeric location index.
In general, when using vertex and fragment functions with TypeGPU pipelines, it is not necessary to set locations on the IO struct properties.
The library automatically matches up the corresponding members (by their names) and assigns common locations to them.
When a custom location is provided by the user (via the d.location attribute function) it is respected by the automatic assignment procedure,
as long as there is no conflict between vertex and fragment location values.
The createGuardedComputePipeline method streamlines running simple computations on the GPU.
Instead of dispatching workgroups, the guarded pipeline allows calling an exact number of GPU threads. Think of it as a parallelized for loop.
Under the hood, it creates a compute pipeline that calls the provided callback only if the current thread ID is within the requested range.
Creates a compute pipeline that executes the given callback in an exact number of threads.
This is different from createComputePipeline() in that it does a bounds check on the
thread id, where as regular pipelines do not and work in units of workgroups.
@param ― callback A function converted to WGSL and executed on the GPU.
It can accept up to 3 parameters (x, y, z) which correspond to the global invocation ID
of the executing thread.
@example
If no parameters are provided, the callback will be executed once, in a single thread.
const fooPipeline = root
.createGuardedComputePipeline(()=> {
'use gpu';
console.log('Hello, GPU!');
});
fooPipeline.dispatchThreads();
// [GPU] Hello, GPU!
@example
One parameter means n-threads will be executed in parallel.
Dispatches the pipeline.
Unlike TgpuComputePipeline.dispatchWorkgroups(), this method takes in the
number of threads to run in each dimension.
Under the hood, the number of expected threads is sent as a uniform, and
"guarded" by a bounds check.
dispatchThreads(5);
// the command encoder will queue the read after `doubleUpPipeline`
var console:Console
The console module provides a simple debugging console that is similar to the
JavaScript console mechanism provided by web browsers.
The module exports two specific components:
A Console class with methods such as console.log(), console.error() and console.warn() that can be used to write to any Node.js stream.
A global console instance configured to write to process.stdout and
process.stderr. The global console can be used without importing the node:console module.
Warning: The global console object's methods are neither consistently
synchronous like the browser APIs they resemble, nor are they consistently
asynchronous like all other Node.js streams. See the note on process I/O for
more information.
Example using the global console:
console.log('hello world');
// Prints: hello world, to stdout
console.log('hello %s', 'world');
// Prints: hello world, to stdout
console.error(newError('Whoops, something bad happened'));
// Prints error message and stack trace to stderr:
// Error: Whoops, something bad happened
// at [eval]:5:15
// at Script.runInThisContext (node:vm:132:18)
// at Object.runInThisContext (node:vm:309:38)
// at node:internal/process/execution:77:19
// at [eval]-wrapper:6:22
// at evalScript (node:internal/process/execution:76:60)
// at node:internal/main/eval_string:23:3
const name = 'Will Robinson';
console.warn(`Danger ${name}! Danger!`);
// Prints: Danger Will Robinson! Danger!, to stderr
Example using the Console class:
const out = getStreamSomehow();
const err = getStreamSomehow();
const myConsole = newconsole.Console(out, err);
myConsole.log('hello world');
// Prints: hello world, to out
myConsole.log('hello %s', 'world');
// Prints: hello world, to out
myConsole.error(newError('Whoops, something bad happened'));
// Prints: [Error: Whoops, something bad happened], to err
Prints to stdout with newline. Multiple arguments can be passed, with the
first used as the primary message and all additional used as substitution
values similar to printf(3)
(the arguments are all passed to util.format()).
The callback can have up to three arguments (dimensions).
createGuardedComputePipeline can simplify writing a pipeline helping reduce serialization overhead when initializing buffers with data.
Buffer initialization commonly uses random number generators.
For that, you can use the @typegpu/noise library.
Creates a compute pipeline that executes the given callback in an exact number of threads.
This is different from createComputePipeline() in that it does a bounds check on the
thread id, where as regular pipelines do not and work in units of workgroups.
@param ― callback A function converted to WGSL and executed on the GPU.
It can accept up to 3 parameters (x, y, z) which correspond to the global invocation ID
of the executing thread.
@example
If no parameters are provided, the callback will be executed once, in a single thread.
const fooPipeline = root
.createGuardedComputePipeline(()=> {
'use gpu';
console.log('Hello, GPU!');
});
fooPipeline.dispatchThreads();
// [GPU] Hello, GPU!
@example
One parameter means n-threads will be executed in parallel.
const fooPipeline = root
.createGuardedComputePipeline((x)=> {
'use gpu';
if (x %16===0) {
// Logging every 16th thread
console.log('I am the', x, 'thread');
}
});
// executing 512 threads
fooPipeline.dispatchThreads(512);
// [GPU] I am the 256 thread
// [GPU] I am the 272 thread
// ... (30 hidden logs)
// [GPU] I am the 16 thread
// [GPU] I am the 240 thread
createGuardedComputePipeline((
x: number
x,
y: number
y)=> {
'use gpu';
const randf: {
seed:typeofrandSeed;
seed2:typeofrandSeed2;
seed3:typeofrandSeed3;
seed4:typeofrandSeed4;
sample:typeofrandFloat01;
sampleExclusive:typeofrandUniformExclusive;
normal:typeofrandNormal;
exponential:typeofrandExponential;
cauchy:typeofrandCauchy;
bernoulli:typeofrandBernoulli;
...7 more ...;
onUnitSphere:typeofrandOnUnitSphere;
}
randf.
seed2: (seed: d.v2f)=>void
Threads do not share the generator's State.
As a result, unless you change the seed in each thread,
each thread will produce the same sequence.
randf.randSeed2 sets the private seed of the thread.
@param ― seed seed value to set. For the best results, all elements should be in [-1000, 1000] range.
seed2(
import d
d.
const vec2f: d.Vec2f
(x:number, y:number)=> d.v2f (+3 overloads)
vec2f(
x: number
x,
y: number
y).
vecInfixNotation<v2f>.div(other: number | d.v2f): d.v2f
Dispatches the pipeline.
Unlike TgpuComputePipeline.dispatchWorkgroups(), this method takes in the
number of threads to run in each dimension.
Under the hood, the number of expected threads is sent as a uniform, and
"guarded" by a bounds check.
dispatchThreads(1024, 512);
// callback will be called for x in range 0..1023 and y in range 0..511
// (optional) read values in JS
var console:Console
The console module provides a simple debugging console that is similar to the
JavaScript console mechanism provided by web browsers.
The module exports two specific components:
A Console class with methods such as console.log(), console.error() and console.warn() that can be used to write to any Node.js stream.
A global console instance configured to write to process.stdout and
process.stderr. The global console can be used without importing the node:console module.
Warning: The global console object's methods are neither consistently
synchronous like the browser APIs they resemble, nor are they consistently
asynchronous like all other Node.js streams. See the note on process I/O for
more information.
Example using the global console:
console.log('hello world');
// Prints: hello world, to stdout
console.log('hello %s', 'world');
// Prints: hello world, to stdout
console.error(newError('Whoops, something bad happened'));
// Prints error message and stack trace to stderr:
// Error: Whoops, something bad happened
// at [eval]:5:15
// at Script.runInThisContext (node:vm:132:18)
// at Object.runInThisContext (node:vm:309:38)
// at node:internal/process/execution:77:19
// at [eval]-wrapper:6:22
// at evalScript (node:internal/process/execution:76:60)
// at node:internal/main/eval_string:23:3
const name = 'Will Robinson';
console.warn(`Danger ${name}! Danger!`);
// Prints: Danger Will Robinson! Danger!, to stderr
Example using the Console class:
const out = getStreamSomehow();
const err = getStreamSomehow();
const myConsole = newconsole.Console(out, err);
myConsole.log('hello world');
// Prints: hello world, to out
myConsole.log('hello %s', 'world');
// Prints: hello world, to out
myConsole.error(newError('Whoops, something bad happened'));
// Prints: [Error: Whoops, something bad happened], to err
Prints to stdout with newline. Multiple arguments can be passed, with the
first used as the primary message and all additional used as substitution
values similar to printf(3)
(the arguments are all passed to util.format()).
Pipeline initialization involves resolving the pipeline code, creating the shader module, and creating the underlying WebGPU pipeline.
This happens automatically the first time the pipeline is executed via .draw, .dispatchWorkgroups, or a similar method.
The initSync method lets you start the initialization early.
To wait until initialization actually finishes on the device (fully avoiding a stall on first execution), use initAsync instead.
Render pipelines require specifying a color attachment for each target.
The attachments are specified in the same way as in the WebGPU API (but accept both TypeGPU resources and regular WebGPU ones). However, similar to the targets argument, multiple targets need to be passed in as a record, with each target identified by name.
Similarly, when using withDepthStencil it is necessary to pass in a depth stencil attachment, via the withDepthStencilAttachment method.
Before executing pipelines, it is necessary to bind all of the utilized resources, like bind groups, vertex buffers and slots. It is done using the with method. It accepts either a bind group (render and compute pipelines) or a vertex layout and a vertex buffer (render pipelines only).
Pipelines also expose a pipe method, which applies a function to the pipeline and returns its result. It lets a package hand out a reusable configuration step that binds everything it owns, instead of documenting a list of with calls.
Pipelines also expose the withPerformanceCallback and withTimestampWrites methods for timing the execution time on the GPU.
For more info about them, refer to the Timing Your Pipelines guide.
After creating the render pipeline and setting all of the attachments, it can be put to use by calling the draw method.
It accepts the number of vertices and optionally the instance count, first vertex index and first instance index.
After calling the method, the shader is set for execution immediately.
Compute pipelines are executed using the dispatchWorkgroups method, which accepts the number of workgroups in each dimension.
The drawIndexed is analogous to draw, but takes advantage of index buffer to explicitly map vertex data onto primitives. When using an index buffer, you don’t need to list every vertex for every primitive explicitly. Instead, you provide a list of unique vertices in a vertex buffer. Then, the index buffer defines how these vertices are connected to form primitives.
Indirect methods read their execution parameters from a buffer, which allows a previous GPU operation to determine the amount of later work.
Mark a typed buffer with .$usage('indirect') and use one of:
Creates a struct schema that can be used to construct GPU buffers.
Ensures proper alignment and padding of properties (as opposed to a d.unstruct schema).
The order of members matches the passed in properties object.
Dispatches compute workgroups using parameters read from a buffer.
The buffer must contain 3 consecutive u32 values (x, y, z workgroup counts).
To get the correct offset within complex data structures, use d.memoryLayoutOf(...).
@param ― indirectBuffer - Buffer marked with 'indirect' usage containing dispatch parameters or raw GPUBuffer
@param ― start - PrimitiveOffsetInfo pointing to the first dispatch parameter. If not provided, starts at offset 0. To obtain safe offsets, use d.memoryLayoutOf(...).
dispatchWorkgroupsIndirect(
const dispatchArgs:TgpuBuffer<d.WgslStruct<{
x:d.U32;
y:d.U32;
z:d.U32;
}>> &IndirectFlag
dispatchArgs);
For an argument block nested inside a larger schema, pass the offset returned by d.memoryLayoutOf:
const
const FrameCommands: d.WgslStruct<{
frameIndex:d.U32;
dispatch:d.WgslStruct<{
x:d.U32;
y:d.U32;
z:d.U32;
}>;
}>
FrameCommands =
import d
d.
struct<{
frameIndex:d.U32;
dispatch:d.WgslStruct<{
x:d.U32;
y:d.U32;
z:d.U32;
}>;
}>(props: {
frameIndex: d.U32;
dispatch: d.WgslStruct<{
x:d.U32;
y:d.U32;
z:d.U32;
}>;
}): d.WgslStruct<{
frameIndex:d.U32;
dispatch:d.WgslStruct<{
x:d.U32;
y:d.U32;
z:d.U32;
}>;
}>
export struct
Creates a struct schema that can be used to construct GPU buffers.
Ensures proper alignment and padding of properties (as opposed to a d.unstruct schema).
The order of members matches the passed in properties object.
Dispatches compute workgroups using parameters read from a buffer.
The buffer must contain 3 consecutive u32 values (x, y, z workgroup counts).
To get the correct offset within complex data structures, use d.memoryLayoutOf(...).
@param ― indirectBuffer - Buffer marked with 'indirect' usage containing dispatch parameters or raw GPUBuffer
@param ― start - PrimitiveOffsetInfo pointing to the first dispatch parameter. If not provided, starts at offset 0. To obtain safe offsets, use d.memoryLayoutOf(...).
dispatchWorkgroupsIndirect(
const commands:TgpuBuffer<d.WgslStruct<{
frameIndex:d.U32;
dispatch:d.WgslStruct<{
x:d.U32;
y:d.U32;
z:d.U32;
}>;
}>> &IndirectFlag
commands,
const dispatchOffset:PrimitiveOffsetInfo
dispatchOffset);
Raw GPUBuffer objects are also accepted.
Offsets must be aligned to four bytes, and the required values must fit in a contiguous region of the buffer.
drawIndexedIndirect additionally requires an index buffer bound with .withIndexBuffer(...).
Render and compute pipelines can record commands into encoders that your application manages.
Pass an existing encoder to .with(...) before drawing or dispatching:
The beginComputePass() method of the GPUCommandEncoder interface starts encoding a compute pass, returning a GPUComputePassEncoder that can be used to control computation.
Completes recording of the render pass commands sequence.
end();
When passed a GPUComputePassEncoder or GPURenderPassEncoder, TypeGPU applies its pipeline and bind-group state to that pass without ending it.
Render pipelines also accept a GPURenderBundleEncoder, allowing .draw(...), .drawIndexed(...), and their indirect variants to be recorded in a render bundle.
The caller remains responsible for ending the pass or finishing the bundle.
When a pipeline is executed directly via draw or dispatchWorkgroups, it records its own pass and submits it to the GPU queue immediately.
For scenarios that require more control, such as executing multiple pipelines in a single render pass or batching multiple passes into a single submission, TypeGPU provides a typed equivalent of the WebGPU command encoder.
It can be created with the createCommandEncoder method on the root object and mirrors GPUCommandEncoder, while accepting TypeGPU resources directly.
Finishes the recording and submits the resulting command buffer to the device queue
submit();
The beginRenderPass method accepts a descriptor similar to WebGPU’s GPURenderPassDescriptor, with a few conveniences:
Attachment views can be TypeGPU textures, texture views and canvas contexts, as well as raw GPUTextureViews.
loadOp, storeOp and depthClearValue default to 'clear', 'store' and 1 respectively. A single color attachment does not need to be wrapped in an array.
occlusionQuerySet and timestampWrites accept TypeGPU query sets as well as raw GPUQuerySets.
There are two equivalent ways to execute pipelines in a pass.
Passing the pass to pipeline.with(pass) keeps the pipeline-centric API, together with all of its with* methods.
Alternatively, the pass itself mirrors the GPURenderPassEncoder API, while accepting TypeGPU resources.
In both cases, the pipeline, bind groups, vertex and index buffers, and the stencil reference are applied lazily when a draw call is recorded, and only if they changed since the previous one.
Both styles operate on the same pass state, resolved per draw with a fixed precedence: state set on the pass (pass.set*) wins over state held by the pipeline (pipeline.with*), which wins over the WebGPU defaults.
Pass state persists until overwritten, regardless of which pipeline is drawn in between, while pipeline-held state applies only to draws of that pipeline.
Note that pipeline.with(pass) allocates a new pipeline wrapper on every call. When drawing the same pipeline repeatedly into one pass, bind it once (const bound = pipeline.with(pass)) and draw with bound in the loop, so that unchanged state is not re-applied on every draw. For values that change between draws, prefer pass-level state (pass.setBindGroup, pass.setImmediates) over the allocating with* methods.
Finishes the recording and submits the resulting command buffer to the device queue
submit();
Calling encoder.submit() finishes the recording and submits it to the device queue.
Shader console.log output and performance callbacks are processed as part of that submission.
Guarded compute pipelines (createGuardedComputePipeline) cannot record into passes or encoders, every dispatchThreads call submits on its own.
Whenever something is not covered by the typed API, the underlying WebGPU resources remain accessible:
root.unwrap(encoder) and root.unwrap(pass) return the raw GPUCommandEncoder, GPURenderPassEncoder or GPUComputePassEncoder, which can be used e.g. for texture copies. Commands recorded this way are invisible to TypeGPU, so after unwrapping a pass, every draw applies its full state again.
encoder.finish() returns the raw GPUCommandBuffer without submitting it, allowing manual batching via device.queue.submit([...]). TypeGPU never sees such a submission, so shader logs and performance callbacks are not processed for it.
Raw GPUCommandEncoders and pass encoders can also be passed to pipeline.with(...) directly, with the same limitations, since TypeGPU cannot know when they are submitted, nor what state has been set on them.
It is also possible to access the underlying WebGPU resources for the TypeGPU pipelines, by calling root.unwrap(pipeline).
That way, they can be used with a regular WebGPU API, though this also requires unwrapping all the necessary resources.
Immediates are small pieces of data set directly on a pass, on a per-draw (or per-dispatch) basis, without the overhead of a buffer. They are WebGPU’s take on push constants.
An immediate variable is defined with tgpu['~unstable'].immediateVar and used in shaders through its .$ property, like other TypeGPU resources.
The value is provided at execution time, either on the pipeline (pipeline.with(immediate, value)) or on the pass (pass.setImmediates), and can be changed between draws.
A value set with pass.setImmediates takes precedence over the pipeline’s value, which takes precedence over the variable’s default. Pass overrides persist across draws and pipeline switches until another setImmediates call replaces them.
Drawing with no value available throws a MissingImmediatesError.
Values are captured (serialized) the moment they are provided, which makes subsequent draws that reuse
them essentially free. Mutating a vector after passing it has no effect until it is set again.
For per-draw changing data, an ArrayBuffer or typed array can be passed instead of a structured
value. The bytes are copied verbatim with no serialization (the caller guarantees they match the
schema’s memory layout), making a thousand different transforms across a thousand draws nearly as
cheap as the raw WebGPU calls.
Immediates come with a few WebGPU-imposed restrictions:
Only one immediate variable can be used in a single shader. To pass multiple values, use a struct schema.
The schema cannot contain arrays, atomics or booleans (use d.u32 or d.i32 instead).
The size of the schema must be a multiple of 4 bytes (pad with d.size if needed, e.g. d.struct({ x: d.size(4, d.f16) })) and is limited by the device’s maxImmediateSize limit (64 bytes by default where supported). A higher limit can be requested at initialization via tgpu.init({ device: { requiredLimits: { maxImmediateSize: 128 } } }).
Immediates require the immediate_address_space WGSL language extension, which is not yet supported everywhere.
Support can be checked via root.enabledWgslLanguageFeatures, and since immediate variables can fulfill accessors (alongside buffer usages and functions), an accessor can seamlessly fall back to a uniform buffer on unsupported hardware.
Retrieves a read-only list of WGSL language extensions supported in the
current environment (navigator.gpu.wgslLanguageFeatures).
Returns an empty set when WebGPU is unavailable.
Defines a variable in the 'immediate' address space, set from the CPU on a per-draw
(or per-dispatch) basis without the overhead of a buffer.
Only one immediate variable can be used in a single shader, and its size is limited
by the device's maxImmediateSize limit. Requires the immediate_address_space
WGSL language extension, check root.enabledWgslLanguageFeatures for support.
@param ― dataType The schema of the held data's type. Cannot contain arrays, atomics or booleans.
@param ― defaultValue The value used when no override is provided via pipeline.with(immediate, value)
or pass.setImmediates. Captured (serialized) at creation time.
Allocates memory on the GPU, allows passing data between host and shader.
Read-only on the GPU, optimized for small data. For a general-purpose buffer,
use
TgpuRoot.createBuffer
.
@param ― typeSchema The type of data that this buffer will hold.
@param ― initial Either initial value of the buffer, or an initializer to execute on the mapped buffer. (optional)
Provides a value for the given immediate variable, used by subsequent draw calls.
The value is captured (copied) at call time; mutating it afterwards has no
effect until it is set again. Takes precedence over values held by the
pipeline (pipeline.with(immediate, value)) and persists across pipeline
switches until set again, like any other pass-level state.
Passing an ArrayBuffer or typed array skips serialization entirely; the bytes
are copied verbatim and the caller guarantees they match the schema's layout.
On the fallback path, buffer.write(...) queues a buffer update rather than recording a command between draws. To use different values for draws within one pass, give each distinct value its own uniform buffer and bind group.