The Simulator Under Your Sink
Two waves are breaking over physical-world AI at the same time. One makes checking your work nearly free. The other lets a model stop describing your faucet and start simulating the fix. Here's where that takes an honest machine.
Today's honest machine works, but it lives inside two hard walls. It can only afford to check its own answer so many times before you'd get bored waiting. And the models it runs can recognize your broken thing, but they can't yet simulate it. Both walls are about to fall, for the same reason: the underlying costs are collapsing. Here's what happens on the other side.
Accuracy isn't one model call. It's many.
People assume a smarter product means a smarter single model. In a safety-critical physical app, it's the opposite. Accuracy is defense in depth: ground the part, score whether the picture is faithful, check whether the step is dangerous, rerank the options, abstain when unsure. Every one of those is another pass, another slice of latency, another fraction of a cent.
So today there's a brutal tradeoff baked into the physics of it: every extra check makes the app slower and pricier. Honesty has a compute bill, and that bill caps how honest you can afford to be inside the few seconds a user will wait.
Speed doesn't just feel nice. It buys accuracy.
This is where the first wave comes in, and it's the one most people read as a UX detail. It isn't. When a class of inference hardware, wafer-scale engines like Cerebras and the broader cohort chasing the same curve, makes a single pass many times faster and cheaper, it doesn't just make the same app snappier. It collapses the tradeoff.
Spend that freed-up budget the obvious way and you get a faster spinner. Spend it the interesting way and you get a more honest machine: the same few seconds now fits an ensemble of critics instead of one, a second and third faithfulness pass, more places the system is allowed to stop and say "I'm not sure." Speed is not the point. Speed is what you spend on depth.
Speed is not a nicety here. It is the enabling condition for depth. Cheaper passes buy more honesty inside the same wait.The speed-to-accuracy paradigm
And once a pass is nearly free, the shape of the product can change too. Not one photo and an eight-second answer, but a live view of the repair as you do it, catching a wrong move mid-motion, the way a patient expert kneeling beside you would. Real-time honesty, not a one-shot guess.
From recognizing your faucet to simulating it
The second wave is bigger. Today's vision models recognize and describe. The hardest unsolved problem in a visual repair app is showing you your exact part being fixed, not a stock clip of someone else's, and it's hard precisely because recognition isn't simulation.
World models, the lineage of systems that learn how the physical world moves and can render plausible futures, change that. Conditioned on your photo, a world model can generate a truthful, physically-grounded sequence of your repair: this cartridge, turned this way, this result. The thing we approximate today with retrieval and a careful downgrade ladder becomes a first-class, personalized simulation.
a worn cartridge"
coming out, correctly
There's a safety dividend hiding in this too. A model that can simulate physics can run the counterfactual: loosen this nut and the line above it drops its pressure and floods the cabinet. Predicting the consequence before a beginner's hand moves is a new layer of protection that recognition alone can never provide. The machine stops being a describer and becomes something closer to a spotter.
Two waves, one machine
These don't arrive separately, and they don't just add. Cheap verification (wave one) and cheap simulation (wave two) compound. A world model that can simulate your repair is only trustworthy if you can afford to verify the simulation, faithfulness, safety, physical plausibility, before it reaches a beginner. Fast silicon is exactly what makes that verification affordable. Each wave is what makes the other safe to ship.
Put them together and the honest machine grows up: from "diagnose your photo and retrieve the best available picture" to "simulate your specific repair, physically, verify it against reality, and only then guide your hands."
Why the data still wins
Here's the part that keeps me calm about all of it. Fast silicon and capable world models are commodities. Everyone will rent the same chips and the same foundation models. Neither is a moat.
What turns a general-purpose world model into one that is actually right about your faucet, that knows this brand's cartridge seats a quarter-turn differently, that a render of bare hands near shards is unsafe, that this fix truly worked, is the one asset nobody else has: the labeled outcome corpus. Every honest abstention, every accepted image, every finished-or-failed repair is a training example the commodity models can't buy. The waves raise the ceiling for everyone. The data is what decides who reaches it, and it's compounding while the hardware catches up.
The waves are coming for everyone. The one who has been quietly writing down every outcome is the one they lift highest.Where the moat actually lives
The crystal ball
So here's the picture a few years out. A complete beginner points a phone at a broken thing. In the time it takes to steady their hand, the machine has simulated the fix on that specific object, checked six ways that the simulation is faithful and that the repair won't hurt them, and started guiding their hands in real time, correcting a wrong move before it becomes a flood or a shock.
And the only score, still, is the one that can't be faked: whether the thing in front of them actually works again when they're done. Bytes to atoms, at the speed of thought, and honest the whole way down.
The future's being built in the open. Come watch.
The simulator isn't here yet. The honest machine that will become it already is. Snap a photo and see.
Install the free alpha →iPhone via TestFlight · no invite code needed