It is a working model of an inference business โ a fifteen-tile dashboard, a live engine map and a control column, all of which have to be on screen at once to tell you anything. There is no honest way to fold that onto a phone, so we would rather show you nothing than show you a broken version of it.
Come back on a laptop and it will be here.
The writeup it is built on reads fine anywhere: read how we did it.
Three ways out. Absorb it and keep going. Relocate to a cheaper grid: power gets cheap again. The cheap grid is a coal-heavy one, so the same tokens carry more carbon; that shows up in EMISSIONS, and some customers read it. Or ask MakerMaker to find the next efficiency improvement now, so you get more tokens per joule out of the same power.
The story ends here for now. Demand keeps rising if you keep serving.
๐คฌ Yeah, no. Use another name. ๐คฌ
MakerMaker can walk you through the board as things come up: what each number means, where a new control appears, which room is the bottleneck. You can skip it later from any of its cards.
You are running an inference business: prompts arrive, your machine answers them, and you are paid per token. The demand for this model is thousands of times bigger than the one GPU you have, and the rent runs every hour whether or not it is busy, so you are losing money from hour zero.
Three things can fix that, and you will need all three: move the price to where the curve peaks, make the machine faster by installing improvements that real engineers really shipped, and once that is paying for itself, grow the fleet to take on more of the demand that is already there. The numbers in here are the published ones.
You lose when the money runs out. Customers turned away or kept waiting are not an instant loss, but they are refunded and they stop coming, so the money goes anyway if you let it run.
Desktop and laptop browsers only for now; mobile is not supported yet.