Product managers and senior leaders are increasingly opening an AI tool, describing a screen or a flow, and getting something polished back within minutes. That output is then handed to a design or engineering team as a starting direction, and in some cases treated as close to final. It looks like a time saving, though the risk it creates gets pushed downstream to whoever builds it next.
Why the output is more persuasive than it should be
Nielsen Norman Group ran a controlled evaluation of this exact scenario, testing AI prototyping tools against a real brief from their own site and comparing the results with what an experienced designer produced for the same page. They tested across three categories of tools and four levels of prompt detail, from a broad one-line brief through to a prompt supported by sketches and a Figma frame. Their conclusion was that these tools can follow an instruction well enough to produce something that looks workable, but they cannot weigh a design trade-off the way a designer with context and judgement can.
That distinction rarely shows up in the output itself. A generated screen that looks finished does not invite the same scrutiny as a rough sketch or a wireframe, even when the thinking behind it is thinner. NN/G found a parallel problem in AI research tools: platforms built without a solid grounding in research methodology produce flawed work presented with confidence, and they do it at scale. The mechanism in prototyping is the same. A solution can be wrong and still look entirely credible.
What gets lost when discovery is skipped
Jumping straight to a solution produces something that looks like design work, but the activity underneath it is different, even where the output ends up looking similar. Discovery and ideation exist to surface things that a fast, prompted output cannot, including:
- The actual problem, framed from evidence rather than from whoever's assumption happened to write the prompt
- The constraints the solution genuinely has to work within
- A reasonable set of alternative directions, so the one chosen was actually chosen rather than simply the first plausible result
- Evidence that could disconfirm the direction, not only evidence that supports it
NN/G's own guidance points the same way: begin from the problem rather than from what a tool happens to be capable of generating. Working backwards from a tool's capability towards a use case tends to produce something disconnected from what people actually need. A prompt only reflects the real problem when whoever wrote it already understood that problem properly, and if that understanding already existed, there would be nothing left to skip.
Why 'it's just a tool' misses the point
The usual defence is that AI is just a tool, and the tool itself isn't at fault. That's true enough, though it misses what's actually being argued here.
NN/G's own testing found that even carefully written, well-specified prompts to AI prototyping tools still needed a fair amount of correction and guidance from an experienced designer before the result was usable. A vaguer prompt performed worse again, producing results that were inconsistent from one attempt to the next. A prompt written by someone without design training, and without any discovery behind it, is likely to sit toward the weaker end of that range by default, not the stronger one. NN/G has also warned specifically that over-relying on AI output without questioning it risks carrying blind spots straight through into the product, and that the way to guard against this is to check AI-generated work against real user research rather than accept it because it looks plausible.
What's actually being criticised here is treating any fast output, whether AI-generated or not, as a substitute for having validated the problem first. AI just makes that mistake easier to make, and easier to make with more apparent confidence than it deserves.
Where this impacts the process
It helps to be specific about what gets skipped, rather than talking about discovery as one vague, undifferentiated stage. Each of these areas fails in a slightly different way, and they tend to compound.
| Risk area | What tends to get skipped | What it costs you later |
|---|---|---|
| Problem framing | Checking whether the stated problem is the real one, rather than a symptom of something else | Building a well-executed solution to the wrong problem, which still has to be unbuilt |
| User needs and context | Talking to anyone who actually does the task the AI-generated screen is designed around | A workflow that fits how the requester imagines the task works, not how it actually works |
| Constraints | Surfacing technical debt, and the regulatory or business rules that shape what's actually feasible | A direction that looks clean in isolation and falls apart the moment it meets the existing system |
| Trade-off reasoning | Weighing one direction against a genuine alternative, rather than accepting the first plausible output | A decision nobody can defend later, because no alternative was ever seriously considered |
| Edge cases and accessibility | Working through what happens for the user who isn't the default case the AI tool assumed | Rework that lands after launch, when it's harder and more expensive to fix |
| Organisational trust | The gradual credibility a design or engineering team builds by being consulted before decisions are made, not after | Teams quietly disengaging from a process that treats their judgement as a formality |
Any one of these on its own is manageable. In practice they tend to show up together, because they all come from the same root cause: a decision made before anyone checked whether it was the right one.
What this misses about product design as a discipline
The deeper issue with this pattern is that it treats product design as though its output is the screen, when the screen is really just what's left over once the actual work has happened: a structured process of reducing uncertainty about what to build before committing resources to building it.
That's what discovery and ideation are for. Producing the interface is downstream work; discovery and ideation are where the decisions actually get tested and defended, and the interface is just the visible residue of that.
NN/G's own recent writing on this makes a related point about experienced designers who appear to skip process and move straight to intuition. What looks like skipped process is usually process that's been compressed and internalised through years of doing the work properly the first few hundred times. An experienced designer moving fast is drawing on a large, tested library of prior discovery. A PM prompting an AI tool for the first time on a given problem has no such library to draw on, however confident the output looks.
What this comes down to is whether a decision that shapes what gets built was actually made with evidence behind it. Tooling, and who's allowed to touch it, is beside the point.
Where the cost actually ends up
The person who generated the shortcut rarely bears the cost of it. That falls to the design team who get briefed on a solution nobody validated, and to the engineering team who build it out.
Rework is the obvious consequence. The less obvious one is what NN/G points to specifically when discussing regulated and high-stakes work, finance among them: in those contexts, process exists to guard against real harm to real people. Skipping discovery on a low-stakes internal tool is a very different risk to skipping it on something that touches how someone manages their money. There is also a harder problem around walking things back. A rough sketch is easy to discard without anyone feeling like progress has been lost. A polished AI mockup already looks like finished work, so raising concerns about it later can feel, to whoever generated it, like undoing progress rather than correcting course before it becomes expensive.
What should actually happen instead
AI design tools, including tools used for ideation and early concept work, belong with the design team. That's a firmer position than telling PMs to use AI carefully, and it follows from the same logic as everything above: using these tools well depends on already understanding the problem space, and that's exactly what a PM or a leader who hasn't done discovery does not yet have.
A PM prompting an AI design tool is making a design decision, about layout, hierarchy, flow, and priority, without the grounding that makes any of those decisions defensible. Treating the result as a rough draft that gets checked later misses what actually happened, because the decision was made at the point the prompt was written, and a review afterwards doesn't undo it.
The role for a PM sits upstream of design tools entirely: describing the perceived problem, sharing evidence, setting constraints, and then handing that to the design team, who are equipped to explore it, including with AI tools of their own. Ideation using AI still has real value. It just needs to happen within the capability that already understands the problem, which is precisely what design teams are for.
Sources
- Nielsen Norman Group, "Good from Afar, But Far from Good: AI Prototyping in Real Design Contexts", nngroup.com/articles/ai-prototyping/
- Nielsen Norman Group, "Don't Start with AI, Start with the Problem", nngroup.com/videos/dont-start-with-ai/
- Nielsen Norman Group, "Design Process Isn't Dead, It's Compressed", nngroup.com/articles/design-process-isnt-dead/

No Comments.