MAKERMAKER.AI
--------------------------------------------------------------

The inference engine simulator needs a bigger screen.

It is a working model of an inference business โ€” a fifteen-tile dashboard, a live engine map and a control column, all of which have to be on screen at once to tell you anything. There is no honest way to fold that onto a phone, so we would rather show you nothing than show you a broken version of it.

Come back on a laptop and it will be here.

--------------------------------------------------------------

The writeup it is built on reads fine anywhere: read how we did it.

capital
$500
hour 0
time to failure
โ€“
capital รท loss rate, at current speed
profit / hr
$0
margin 0%
cost per 1M tokens
โ€“
against what you charge
efficiency
1.0x
vs factory default
emissions
0
tonnes CO2 so far
market share
โ€“
of the work you take
customers waiting
0
arriving 0/s
revenue / hr
$0
at the list price / 1M tok
costs / hr
$0
rental ยท power ยท refunds
tokens a second
โ€“
answer tokens leaving the machine

GPUs

โ€“
OUTPUT ยท
SERVINGclock: seconds MODELclock: milliseconds โ€” ร—10ยณ faster KERNELSclock: microseconds SILICONclock: nanoseconds โ€” ร—10ยณ again FLEETclock: days โ€” the slow loop that owns the fast ones BEDROCK โ€” our own bench, one MI300X: 1,611 factory default ยท 2,418 on the tuned public recipe INTAKE SCHEDULER KV ALLOC fused bookkeeping โœ“ KV-fusion BATCH B(14) C(30) max batch: 256 max-batch embed attention router experts combine sample โŸฒ AUTOREGRESSIVE STEP โ€” the output token becomes the next input โ†ฉ ร—k tokens / lap ร—36 EMBED ATTENTION ROUTER EXPERTS COMBINE SAMPLER draft ร—k โœ“ EAGLE-3 CARTRIDGE: GPT-OSS-120B โ€” the model this machine serves DISPATCH ENV=โ€ฆ ร—7 torch.compile env-flags ATTN KERNEL AITER GEMM your 3 hand entries launch drawer: 437 shapes masked tail tile tables FUSED OPS fused top-k CU ARRAY MM@2 ATTN โ—‚ โ–ธ MOE IF LINKS 8 replicas โ€” each GPU serves its own requests; one scheduler fans them out fleet ร—8 A2A โ–ธ to the experts' hosts, and home ร—8 โณ barrier: waiting for GPU 3 OP LAUNCH TILES CORES requests ยท s 0 tokens ยท ms 0 launches ยท ยตs 0 FLOPs ยท ns โ€” a blur 0 discoveries ยท days 0
LOG ยท โ€“
TIME LEFTโ€“
#namescore

CONTROLS

TIME ร—1.0

WHAT YOU SELL GPT-OSS-120B

SERVICE LEVEL AGREEMENT
customers wanting itwhat you can serveyour price

Machine down

INCOMING 0 waiting the game is paused while this is open

LOG the game is paused while this is open

What the benchmark turned up

Rented.

Electricity price rises

Three ways out. Absorb it and keep going. Relocate to a cheaper grid: power gets cheap again. The cheap grid is a coal-heavy one, so the same tokens carry more carbon; that shows up in EMISSIONS, and some customers read it. Or ask MakerMaker to find the next efficiency improvement now, so you get more tokens per joule out of the same power.

The machine runs itself.

The story ends here for now. Demand keeps rising if you keep serving.

Leaderboard

#namescore

Time.

#namescore
Share

Want the tutorial this round?

MakerMaker can walk you through the board as things come up: what each number means, where a new control appears, which room is the bottleneck. You can skip it later from any of its cards.

Someone has rented you one GPU. Nobody tuned it.

You are running an inference business: prompts arrive, your machine answers them, and you are paid per token. The demand for this model is thousands of times bigger than the one GPU you have, and the rent runs every hour whether or not it is busy, so you are losing money from hour zero.

Three things can fix that, and you will need all three: move the price to where the curve peaks, make the machine faster by installing improvements that real engineers really shipped, and once that is paying for itself, grow the fleet to take on more of the demand that is already there. The numbers in here are the published ones.

You lose when the money runs out. Customers turned away or kept waiting are not an instant loss, but they are refunded and they stop coming, so the money goes anyway if you let it run.

Desktop and laptop browsers only for now; mobile is not supported yet.