Meta
Hyperscale AI infrastructure
At hyperscale, a few percent of kernel throughput adds up to thousands of GPUs. That is the problem our optimization work exists to solve.
Company
MarsCompute comes out of long-running research on GPU acceleration. We build the optimization layer between models and silicon, and a model API on top of it.
Where this comes from
MarsCompute is led by Prof. Bingsheng He of the National University of Singapore, with a team working on kernels, compilers and AI systems.
A research project on GPU acceleration led by Prof. He, started the year CUDA was born.
Automated optimization results published for FlashAttention-2, 3 and 4.
The intra-kernel tracer in G-Watch goes open source, with a paper describing it.
One OpenAI-compatible endpoint for open and commercial models, from the same team.
Who we work with
Our customers include Meta and 4Paradigm, among others, and we collaborate with AI Singapore.
Meta
At hyperscale, a few percent of kernel throughput adds up to thousands of GPUs. That is the problem our optimization work exists to solve.
4Paradigm
Enterprise AI has to perform on the hardware a business already owns. Getting that across mixed accelerators is a planning problem, not a hand-tuning one.
AI Singapore
AI Singapore develops the SEA-LION family for the region's languages. Qwen-SEA-LION v4.5 is live on our API today.
Work with us
We work with teams that run production AI systems, where performance on the hardware is critical.
A technical look at your workloads and the hardware they run on.
Collaboration with hardware, platform and ecosystem teams.
Work with us directly as a design partner.
Who we build for
Request access to the API, or talk to the team about your own kernels and hardware.