G-Watch · open source

See what a GPU is really doing.

G-Watch is our open-source toolbox for analysing GPU execution. Its tracer, Xtrace, records what happens inside a kernel without changing the kernel.

G-Watch · open source

The toolbox that shows an agent what a GPU is doing.

G-Watch is our analysis framework for GPU execution. It gives engineers, and the coding agents working beside them, exact data to optimize with.

  • Xtrace. A timeline of every warp inside a kernel, traced on the compiled binary.
  • Binary analysis. What the compiler really emitted, down to the instruction.
  • Microbenchmarking. Small, targeted measurements that test an idea before a kernel is changed.
  • Built for agents. Installs as a skill, so a coding agent can profile and tune a kernel on its own.
# the toolbox
pip3 install gwatch

# the skill for your coding agent
npx skills add mars-compute-ai/G-Watch -g

Xtrace

Trace the kernel you ship, not a different one.

A trace is a measurement. It is only useful if the kernel being measured stays the same.

Existing tracers
Source or IR probes added here
Compiler builds the probes in
Binary a changed kernel
Xtrace
Source or IR untouched
Compiler optimizes as usual
Binary probes added here

Existing tracers add probes before the kernel is compiled. The compiler treats them as part of the kernel, schedules the code differently, and some optimizations disappear. Xtrace edits the compiled binary and fits each probe into room the kernel leaves unused. It changes no basic block.

Original instructions still in the traced kernel

Xtrace
94–98%
Existing tracers
8–48%

Time added to the kernel

Xtrace
0.9–2.8%
Existing tracers
3.8–75.6%

Ranges across the kernels tested, on NVIDIA H100 and B300 and AMD MI300X. Source: G-Watch, September 2026.

Closed-source kernels too

Xtrace needs the binary, not the source. That opens up the kernels inside vendor libraries.

cuDNN cuBLAS TensorRT-LLM

19 GPU architectures

Both vendors, from the newest parts back several generations.

NVIDIA Volta to Blackwell AMD CDNA1 to CDNA4 AMD RDNA1 to RDNA4

Six kernel languages

However the kernel was written, the trace is taken the same way.

CUDA HIP CuTe DSL Triton TileLang FlyDSL

Open Traces

Real traces, open in your browser.

We publish Xtrace traces of open and closed-source kernels. Each one opens as an interactive timeline, warp by warp.

Software traced so far

FlashAttention-3 FlashAttention-4 cuDNN cuBLAS TensorRT-LLM FlashInfer FlashMLA HSTU

See what the tools found.

Published gains on FlashAttention kernels, with the conditions each was measured under.