Totals hide the cause
Hardware counters report totals over a whole kernel launch. They show that a kernel is slow, not which stage is waiting on which.
Technology
More and more, what limits a model in production is not raw compute. It is how well its kernels fit the hardware. We build the tools that measure that fit and the system that improves it, on NVIDIA and AMD.
The problem
A modern accelerator can deliver far more than most workloads get out of it. Three things stand in the way.
Hardware counters report totals over a whole kernel launch. They show that a kernel is slow, not which stage is waiting on which.
NVIDIA, AMD and mixed fleets each need their own execution plan. Tuning done by hand for one architecture does not carry over to the next.
Changes to kernels, memory and runtime affect one another. They have to be planned together, and manual tuning loops do not scale to that.
The stack
Select a layer. The models live today are served through their providers; the layers below are what open models on our own GPUs are tuned with.
Layer 01
Qwen-SEA-LION v4.5 and GPT-6 Astra are live today. Kimi K3 and GLM-5.3 are next, on GPU capacity we operate.
Layer 02
Keys scoped to one model and one member. Cost reserved against your spending limit before a request is sent. Every request metered by token and exportable.
Layer 03
It discovers optimization opportunities from how workloads really execute, models how kernels, memory and runtime interact, and plans changes that do not conflict with each other.
Layer 04
G-Watch combines intra-kernel tracing, binary analysis and microbenchmarking. It gives engineers, and the coding agents working beside them, exact data instead of totals.
Layer 05
NVIDIA from Volta to Blackwell, and AMD across CDNA and RDNA. Kernels written in CUDA, HIP, CuTe DSL, Triton, TileLang or FlyDSL, including closed-source vendor libraries.
Mars Optimization Brain
Three coordinated layers turn what the hardware and the workload are doing into changes that are safe to apply.
Maps the topology, memory hierarchy, interconnect behaviour and execution constraints of each target environment.
Builds an optimization graph and generates conflict-aware plans across kernels, runtime policy and scheduling paths.
Applies each optimization, measures its effect, and keeps improving as architectures, drivers and serving patterns change.
What this means for the API
The API and the optimization stack come from one team. This is exactly where they meet today.
Provider-served models
Qwen-SEA-LION v4.5 and GPT-6 Astra run on their providers' own infrastructure. Through MarsCompute you get one API for them, with scoped keys, spending limits and per-request metering.
Models on our GPUs
Kimi K3 and GLM-5.3 will run on GPU capacity we operate. That is where this stack applies in full: we control the kernels, the runtime and the hardware they run on.
Request access to the API, or talk to the optimization team about your own kernels and hardware.