Company

A GPU systems team, building a model platform.

MarsCompute comes out of long-running research on GPU acceleration. We build the optimization layer between models and silicon, and a model API on top of it.

Where this comes from

GPU systems research since 2007.

MarsCompute is led by Prof. Bingsheng He of the National University of Singapore, with a team working on kernels, compilers and AI systems.

  1. 2007

    Mars begins

    A research project on GPU acceleration led by Prof. He, started the year CUDA was born.

  2. March 2026

    First kernel reports

    Automated optimization results published for FlashAttention-2, 3 and 4.

  3. September 2026

    Xtrace released

    The intra-kernel tracer in G-Watch goes open source, with a paper describing it.

  4. Today

    The model API

    One OpenAI-compatible endpoint for open and commercial models, from the same team.

Who we work with

Working with teams that run AI at scale.

Our customers include Meta and 4Paradigm, among others, and we collaborate with AI Singapore.

Customer

Meta

Hyperscale AI infrastructure

At hyperscale, a few percent of kernel throughput adds up to thousands of GPUs. That is the problem our optimization work exists to solve.

Customer

4Paradigm

Enterprise AI platforms

Enterprise AI has to perform on the hardware a business already owns. Getting that across mixed accelerators is a planning problem, not a hand-tuning one.

Collaborator

AI Singapore

Models for Southeast Asia

AI Singapore develops the SEA-LION family for the region's languages. Qwen-SEA-LION v4.5 is live on our API today.

Work with us

Three ways to start.

We work with teams that run production AI systems, where performance on the hardware is critical.

Technical architecture review

A technical look at your workloads and the hardware they run on.

Partnership

Collaboration with hardware, platform and ecosystem teams.

Design partner program

Work with us directly as a design partner.


Who we build for

Teams for whom performance is infrastructure.

  • Hyperscalers operating AI fleets across more than one architecture.
  • AI infrastructure and platform teams inside enterprises.
  • GPU and accelerator vendors scaling production adoption.
  • Model serving teams focused on latency and efficiency.
  • Developers who want models behind one well-run API.

A model to call, or a fleet to tune?

Request access to the API, or talk to the team about your own kernels and hardware.