Index · built in the open

Projects

Benchmarks, platforms and recipes. Measured, not promised, and every figure recorded on the hardware.

01

Inference Atlas

A collaborative benchmark platform for LLM inference engines. Submit a run and get comparable throughput and latency across engines, models, quantisations and hardware — plus a terminal app that writes the serving recipe for you.

TypeScriptvLLMSGLangBenchmarks
02

Agents Arena

An open platform where AI agents compete in games. Any agent that speaks HTTP can join; the arena handles turns, rule enforcement and the leaderboard, so the only thing being compared is the agent.

HTTPAgentsLeaderboard
In development Coming soon
03

NVIDIA DGX Spark

A collection: everything I have published about serving large language models on a single GB10 box — 180B parameters offloaded to NVMe, NVFP4 against AutoRound against FP8, decode strategies worth 7× on untouched weights, and one-command installs. Five projects, all measured on the hardware.

GB10vLLMQuantisation5 projects

More of this ends up on Medium before it ends up here.

Read the writing