Xtrace, our binary-level GPU kernel tracer, is now open source →

Model-as-a-Service, engineered from the silicon up

The model API from the team that tunes the silicon.

We build the tools that see inside GPU kernels and the system that tunes them, then put models behind one OpenAI-compatible API.

  • Kernel-level optimization
  • NVIDIA and AMD
  • OpenAI-compatible API
Attention kernel, per-warp timelineKernel timeline
Producer
MMA
Softmax A
Softmax B
Epilogue
Tiles finished in this window 3
Matrix-multiply warp stalled 48%

A schematic, not a measurement. Real traces of FlashAttention, cuDNN and more are in G-Watch Open Traces.

  • NVIDIA Volta to Blackwell
  • AMD CDNA1 to CDNA4
  • AMD RDNA1 to RDNA4
  • CUDA
  • HIP
  • CuTe DSL
  • Triton
  • TileLang
  • FlyDSL
  • cuDNN
  • cuBLAS
  • TensorRT-LLM

Why MarsCompute

Strengths you can measure.

94–98%

We trace the kernel you ship

Xtrace leaves 94 to 98% of a kernel's instructions intact. Existing tracers keep 8 to 48%.

3.9×

Agents do the tuning

A coding agent needed 3.9× fewer iterations to tune a kernel with our trace.

19

NVIDIA and AMD alike

19 GPU architectures covered, from Volta to Blackwell and CDNA to RDNA.

+50.8%

Gains you can check

Our best published gain so far: FlashAttention-2 on an NVIDIA A100.

Start building on MarsCompute.

Request access and we will set up your workspace. Bring the client you already use.