94–98%
We trace the kernel you ship
Xtrace leaves 94 to 98% of a kernel's instructions intact. Existing tracers keep 8 to 48%.
Model-as-a-Service, engineered from the silicon up
We build the tools that see inside GPU kernels and the system that tunes them, then put models behind one OpenAI-compatible API.
A schematic, not a measurement. Real traces of FlashAttention, cuDNN and more are in G-Watch Open Traces.
What we do
Open and commercial models behind one OpenAI-compatible endpoint, with keys, limits and metering built in.
Explore the platform 02The Mars Optimization Brain plans and applies kernel-level optimizations across NVIDIA and AMD hardware.
See the technology 03Our open-source toolbox that lets engineers and coding agents see inside a GPU kernel.
Meet G-WatchWhy MarsCompute
94–98%
Xtrace leaves 94 to 98% of a kernel's instructions intact. Existing tracers keep 8 to 48%.
3.9×
A coding agent needed 3.9× fewer iterations to tune a kernel with our trace.
19
19 GPU architectures covered, from Volta to Blackwell and CDNA to RDNA.
+50.8%
Our best published gain so far: FlashAttention-2 on an NVIDIA A100.
Model library
Two are live today. Two more are coming on GPUs we operate.
Who we work with
Request access and we will set up your workspace. Bring the client you already use.