You have an idea. You also have, if you're honest, no real idea whether it's a good one yet. That's the actual starting point for MVP development, not conviction, but an unanswered question worth the cost of answering it properly.
Define: what you're actually testing
Before anything gets built, the real job is deciding what the product needs to prove and cutting everything that doesn't serve that one thing. This is the step most often skipped, because it's the least fun one, it feels like a delay when everyone's eager to start building. But a first version scoped around the wrong question is expensive regardless of how well it's built. Getting this step wrong is the single most common reason a first version teaches nobody anything.
In practice, this looks like a short, occasionally uncomfortable exercise: writing down what would have to be true for the idea to work, then identifying which of those things is actually in doubt versus just assumed. Everything that's genuinely assumed, not tested, just believed, gets cut from the first version's scope, even if it's tempting to include. The definition phase ends with a short, specific list: this version exists to find out X, for a person who looks like Y, doing Z. If any part of that sentence is vague, the rest of the process inherits the vagueness.
Build: where speed actually helps, and where it doesn't
This is where agent-assisted building earns its reputation in MVP development, scaffolding, boilerplate, a working first draft, all dramatically faster than it used to be. That speed is real and worth using fully. What doesn't get automated, and shouldn't be, is the handful of decisions that are expensive to get wrong: how the data is modeled, what happens when a dependent service fails, which shortcuts are fine to ship and which will need revisiting in a month.
This is the actual shape of how we build first versions at Hexsis: agent-assisted engineering for everything that benefits from speed, paired with human verification at the decisions that don't. Not review as a formality after the fact, review at the specific points where getting it wrong is expensive, before the shortcut becomes load-bearing.
This is also where scope discipline gets tested for real. It's easy to agree in the definition phase that the first version will be narrow, and then quietly widen it during the build once momentum kicks in and a dozen small "while we're at it" additions each seem harmless on their own. None of them are harmless. Each one adds a piece of surface area that has to be maintained, explained, and eventually reconciled with whatever the real answer to the original question turns out to be.
Launch: in front of real use, not a demo
Launching a first version means putting it in front of people actually trying to do the real task it exists for, not a private beta of friends being polite, not a room of investors watching a guided walkthrough. Feedback from a demo tells you whether the product is convincing. Feedback from real use tells you whether it's true. Only one of those is worth building the next version around.
A launch that's actually useful comes with a plan for what happens to the answer once it arrives, who's watching what, what would count as a clear yes, what would count as a clear no, and what happens if the result is the uncomfortable third option: maybe, depending on something nobody thought to ask about. Without that plan, real usage data tends to get interpreted however whoever's looking at it wants it to be interpreted, which defeats the entire purpose of building something real in the first place.
What good documentation looks like during this process
Not a wiki nobody updates, a short, living record of what the first version is supposed to prove, updated whenever the scope changes, and visible to everyone deciding what goes in or stays out. When someone proposes adding something mid-build, the fastest way to evaluate it isn't a debate about whether it's a good idea in general. It's checking it against that one document: does this help answer the actual question, or not? Most scope arguments resolve themselves the moment they're framed that way, because most additions, honestly assessed, don't.
Where each phase most commonly breaks
Define breaks when the question quietly gets swapped for a feature list somewhere in the second meeting, because a feature list feels more actionable than an open question and nobody notices the substitution happening. Build breaks when scope creep gets rationalised one small addition at a time, each one defensible in isolation, until the "smallest real version" has become something considerably larger without anyone deciding that on purpose. Launch breaks when the team gets nervous about real feedback and quietly narrows who gets to see the product first, filtering it down to people who were always going to be kind about it, which produces confidence instead of information, and confidence is the one thing a first version was never supposed to be optimising for.
Each of these failures looks reasonable the moment they happen. That's what makes them worth naming ahead of time, before the meeting where it feels like the sensible thing to do.
Three questions carry the whole process: what are we trying to find out, what's the smallest real thing that finds it out, and are the people using it the ones who'd actually pay for the answer. Everything else is detailed.
