One PM, One Designer, One Engineer: How AI-Native Teams Build Now
It is Tuesday morning and a product manager is running the company's checkout
page on their laptop. Not a mockup of it. The actual app, on localhost, on a
branch called proto/split-payments, with seed data loaded. They ask Claude
Code to let a customer split an order across two cards. Four minutes later
they are clicking through it. The confirmation screen looks wrong when one
card is a gift card with a low balance, so they call the designer over, and
the two of them fix the layout in the running product while the agent keeps
up.
By lunch they have found six edge cases nobody would have written into a requirements document, killed one idea (three-way splits), and written down why. No engineer has touched it yet. No Figma file exists. The engineer will be pulled in on Thursday, once the PM and designer are satisfied, and their job will not be "build this." It will be "make this safe to ship."
This is what a product team looks like in a company that builds with AI natively: one PM, one designer, one engineer, a set of agents, and roles that overlap so much they are close to interchangeable. The shape works. It is also much harder to set up than it looks, because the whole thing rests on a verification loop most teams do not have. That loop is the problem we are working on at CatalEx, and we will have more to share on it soon. Stay tuned.
This post is the playbook: how the new shape works, the tools teams use at each stage, where humans still read code and where they do not, what breaks, and what to do in each situation.
The old shape was a relay race
For about fifteen years, product work moved through a relay. Each person produced an artifact whose only purpose was to be read by the next person.
- The PM wrote a PRD.
- The designer turned the PRD into mockups.
- The engineer turned the mockups into code, discovering along the way that the PRD skipped half the edge cases.
- QA tested the code against the PRD, which was by now out of date.
Every arrow in that chain is a lossy translation. The PRD cannot tell you that the confirmation screen overflows on a small phone. The mockup cannot tell you the refund API rejects partial amounts. You only learn those things by running the product, and in the relay, the first person who runs the product is the engineer, at the most expensive point to change your mind.
The team shape followed from the relay. You needed several engineers per PM because engineering was the slow, expensive stage where all the discovery happened.
The new shape: discovery moves to the product
Coding agents changed which stage is expensive. Producing a working version of a feature is now cheap. Producing a correct, safe, maintainable version is still expensive. So AI-native teams split the work along that line.
In the relay, discovery happens in the red box, late and in code. In the new loop, discovery happens in the first box, early and in the product. The prototype is no longer a picture of the feature. It is the feature, running, unhardened.
That one change reshapes every role:
| Role | In the relay | In the AI-native team |
|---|---|---|
| PM | Writes the PRD, answers questions | Prototypes on the product, finds edge cases by clicking, owns the decision log |
| Designer | Draws mockups | Shapes look and feel in the running product, guards the design system |
| Engineer | Translates mockups into code | Takes a working prototype to production: failure modes, privacy, security, performance, observability |
| QA | Writes test cases against the PRD | Reviews the spec and its acceptance criteria, owns what "done" means |
| Agents | Autocomplete | Write most of the code, tests, and docs at every stage |
The roles converge because they all now work on the same artifact. The PM and the designer are both editing the running product. The engineer reads the same decision log the PM wrote. A designer who can run Claude Code can prototype without the PM, and a PM with taste can adjust spacing without the designer. What stays distinct is judgement, not tooling: the PM owns what the product should do, the designer owns how it should feel, and the engineer owns whether it holds up.
How the loop actually works
There are five stages. Each has an owner, a place the work happens, and an artifact it leaves behind. The artifact is what makes the next stage possible without a meeting.
1. Explore on a canvas
Before anyone touches the product, the team dumps the problem somewhere visual: user quotes, screenshots of competitors, rough flows, the metric they want to move. Most teams still do this in Miro or FigJam. The difference in 2026 is that the canvas has agents on it. Miro's Sidekicks and Flows can cluster sticky notes, draft user flows from interview transcripts, and pull in context from connected tools, so the brainstorm ends with a structured board instead of a photo of a wall.
Artifact: a board with the problem statement, the flows worth trying, and the questions nobody can answer yet.
2. Prototype on the real product
This is the stage that changed the most. The PM and designer stopped designing in a separate tool and started prototyping in the product itself. They clone the repo, run it locally with seed data, create a branch, and ask a coding agent to build the flow.
Why the real product and not a blank canvas? Because the real product has the real constraints. It has the design system, the existing navigation, the actual data shapes, the loading states. A prototype built there shows you the edge cases a mockup hides: the empty cart, the expired session, the currency with no decimal places.
For visual exploration before touching code, some teams start in Claude Design, which reads a codebase's design system and generates interactive prototypes from conversation, then hand the chosen direction to Claude Code to build into the app. Others go straight to Claude Code, Codex, or Cursor.
The safety of this stage depends on rules the agent follows in prototype mode.
Teams put them in the repository's agent instructions file (CLAUDE.md or
AGENTS.md), so every PM and designer gets the same guardrails without having
to remember them:
Applies to any branch named proto/*.
- Never push to main. Never merge. Prototypes are reviewed, not shipped.
- Put new behaviour behind the flag
proto_<feature>, default off. So a prototype can be demoed on staging without affecting anyone. - Build only with components from
src/ui. If a component is missing, stop and log it indecisions/<feature>.mdinstead of inventing one. Invented components are how a design system dies. - Use
fixtures/for data. Never call production APIs, never read real customer records. Prototypers are not trained to handle production data. - After every change, append one line to
decisions/<feature>.md: what changed and why.
Artifact: a branch with a working, flagged, unhardened feature.
3. Research and log every decision
Once the PM and designer are satisfied with how the feature behaves, they switch from building to documenting. This is where research happens: checking the prototype against interview notes, looking at analytics for how often the edge cases actually occur, and asking an agent to read the diff and list every behaviour the prototype introduced.
The core output is a decision log. Every choice made during prototyping gets a record, including the ideas that were killed. The killed ideas matter most, because six months later someone will propose three-way splits again, and the log is what saves the team a week.
Here is what one record looks like. It fits on an index card, and that is deliberate: a record that takes more than two minutes to write does not get written.
| Field | Entry |
|---|---|
| Decision | Cap split payments at two methods. |
| Decided by | PM and designer |
| Why | Three-way splits doubled the failure states on the confirmation screen, and none of the 14 customer interviews asked for more than two. |
| Rejected | Three methods. The prototype branch is kept as evidence. |
| Open question | Refund to which method when one card has expired? Owner: engineer. |
| Risk tier | High. It touches money, so an engineer must read the final diff. |
Two fields do more work than they appear to. Open question hands unknowns to the engineer by name, so they become hardening work instead of production surprises. Risk tier is read later by the verification loop to decide how much human review the change gets.
From the prototype and the log, the agent drafts a spec: behaviour, constraints, and numbered acceptance criteria. The PM edits it. QA reviews it. The spec is now the source of truth, and the prototype is evidence for it.
Behaviour
- A customer may pay with at most two methods (DEC-0142).
- The first method is charged the amount the customer enters; the second is charged the remainder.
Acceptance criteria
- AC-1: A gift card with a balance below the order total pre-fills as the first method with its full balance.
- AC-2: If the second charge fails, the first charge is voided within 60s and the customer sees the "payment not completed" state.
- AC-3: No card number, full or partial, appears in client logs.
- AC-4: A refund on a split order returns funds to each method in proportion to its charge.
Artifact: a decision log, and a spec with acceptance criteria QA has signed off.
4. Harden: the engineer arrives
Only now does the engineer take ownership. They are not translating anything. They read the spec, the decision log, and the prototype diff, and they ask a different set of questions from the ones the PM asked:
- How does it fail? What happens when the payment provider times out between the two charges? Is AC-2 even achievable with this provider's API?
- What does it leak? Which fields reach logs, analytics, or third-party scripts?
- Who can abuse it? Can a user split an order so the second charge is for a cent, and does that bypass fraud checks?
- Does it scale? What did the agent do that works on seed data and falls over on a customer with 4,000 orders?
- Can we see it? Are there metrics and alerts for the new failure states?
Agents help here too, but in a narrower role. The engineer runs a security review agent over the diff, has an agent enumerate failure paths through the payment code, and asks for tests covering each one. The engineer decides which findings are real.
Often the engineer rewrites parts of the prototype. That is expected. The prototype's job was to be right about what; the engineer's job is to be right about how.
Artifact: a production-ready pull request, behind the same flag.
5. The verification loop
The last stage is not a stage a person performs. It is a loop that runs on every change, and it is what lets this team avoid reading every line of agent-written code.
The piece most teams miss is the spec coverage check. Agents are very good at writing tests that pass. They are less reliable at writing tests that check the thing the spec cares about. So the loop enforces a link: every acceptance criterion in the spec must be named by at least one end-to-end test, and every test must name the criterion it covers. The check itself is a few lines of script: read the criterion ids out of the spec, look for each one in the test titles, and fail the build if any are missing.
A covered criterion is not the same as a well-tested one, so QA also reads what each test does. Here is what a good test for AC-2 walks through, in plain steps:
- Open checkout with a $120 test order.
- Turn on split payment.
- Enter a valid test card as the first method.
- Enter a sandbox card that always declines as the second method.
- Press Pay.
- The screen shows "Payment not completed".
- The payment ledger shows no money held on the first card.
Step 7 is the one agent-written tests tend to skip. A screen can say "voided" while the money is still held on the customer's card. Test the system of record, not just the screen.
The gates run in the order that fails fastest, so a broken change is rejected in seconds rather than after a ten-minute browser run:
| Order | Gate | Fails when |
|---|---|---|
| 1 | Unit tests | Any function-level behaviour breaks |
| 2 | Spec coverage | An acceptance criterion has no test naming it |
| 3 | End-to-end tests | The product does not do what a criterion says |
| 4 | Security and privacy scans | A secret is committed, or card or personal data reaches a log |
| 5 | Review gate | The decision log says high risk and no engineer has approved the diff |
The last gate is where the risk tier from the decision log pays off. A low tier change needs a green loop and a look at the preview environment. A high tier change needs a green loop and an engineer's approval on the diff.
The tools, stage by stage
The tool list moves fast, so read this as a snapshot of what teams are using in the second half of 2026, not an endorsement.
| Stage | Common tools | What they are used for |
|---|---|---|
| Explore | Miro (with Sidekicks and Flows), FigJam | Clustering research, drafting flows, keeping context on a shared board |
| Visual exploration | Claude Design, Figma | Interactive prototypes grounded in the design system; Figma remains the home of the design system itself |
| Prototype on product | Claude Code, Codex (CLI and app), Cursor | Building the feature in the running app on a local branch |
| Research and log | Claude or ChatGPT with product analytics and tracker connectors (Linear, Jira, PostHog via MCP) | Checking edge cases against real usage, drafting decision records and specs |
| Harden | VS Code or Cursor with Claude Code or Codex, security review agents | Reading diffs, running agents in parallel, inspecting their work |
| Verify | Playwright, the unit test runner, GitHub Actions, preview environments, feature flags | The loop that decides whether a change is acceptable |
Does anyone still open VS Code?
Yes, and the reason is instructive. The engineers still live in an editor, but they use it differently. They are not typing most of the code. They are running several agents at once, watching what each one does, reading diffs, and stepping in when an agent goes somewhere it should not. The editor is a supervision console.
The reason to keep one is the user. When a change is headed to customers and touches money, data, or permissions, someone accountable has to see the code. An editor with a good diff view is still the fastest way to do that.
And when does nobody look at the code?
When the risk is low and the loop is trustworthy. A copy change, a new empty state, a tweak to an internal admin screen: for these, teams look at the output on the product itself, in a preview environment, and trust the loop for the rest. Reading the diff of a tooltip change adds cost and catches nothing the e2e test would not.
The mistake is not skipping code review. It is skipping code review without a loop that earns it. "We don't read the code" is a reasonable policy when the spec is reviewed, the acceptance criteria have tests, and the tier is low. It is a dangerous one when any of those is missing.
What goes wrong
The shape fails in a small number of recognisable ways. Each has a check you can automate or make a habit.
| Failure mode | What it looks like | The check that catches it |
|---|---|---|
| Prototype ships by accident | A proto/* branch gets merged because the demo looked finished | Branch protection rejects proto/* merges; hardening happens on a new branch |
| Decisions live in chat | The engineer asks "why two cards?" and nobody remembers | PR template requires a decision log link; CI fails if a flagged feature has no decisions/ folder |
| Spec drifts from the prototype | The spec says four behaviours, the prototype has seven | An agent diffs prototype behaviour against the spec before QA review, and lists the extras |
| Verification theatre | Agent-written tests pass but assert nothing the spec cares about | Spec coverage check, plus QA reviews test titles against AC ids |
| Hardening skipped under pressure | "It works on staging, ship it" | The risk tier gates merge; high tier requires engineer approval |
| Design system erosion | Each prototype adds a slightly different button | Prototype rules forbid new components; lint fails on UI imports outside src/ui |
Why the obvious fix fails
The obvious move is to give every PM and designer Cursor or Claude Code and declare the team AI-native. Teams that do only this get a burst of excitement and then a mess.
Prototypes pile up as branches nobody can evaluate. Each one is a few thousand lines of agent-written code with no spec, no decision log, and no tests. The engineer is asked to "just clean it up," which is harder than building it from scratch, because they first have to reverse-engineer what the prototype was supposed to do. Engineering becomes the cleanup crew, the engineer burns out, and leadership concludes that PMs should not prototype.
The tools were never the hard part. The hard part is the connective tissue: the prototype rules, the decision log, the spec with numbered criteria, the coverage check, and the risk tiers. Without that, overlapping roles do not make a team faster. They make ownership unclear.
A worked example: split payments, before and after
The same feature, built by the same company, a year apart.
| Step | Relay (2025) | AI-native loop (2026) | Why it matters |
|---|---|---|---|
| Define | PM writes an 8-page PRD over a week | PM and designer prototype on the product in two days | Edge cases are found by clicking, not guessed |
| Design | Designer mocks 12 screens in Figma | Designer fixes layouts in the running app | The gift card overflow is caught on day one |
| Decide | Decisions scattered across comments | 9 records in decisions/split-payments/, 3 rejected options | Three-way split is not re-litigated next quarter |
| Spec | PRD goes stale during build | Spec with AC-1 to AC-4 extracted from prototype, QA signs off | QA owns "done" before code is hardened |
| Build | Two engineers, three sprints | One engineer hardens in four days | Engineer time goes to failure, privacy, abuse |
| Verify | Manual QA pass before release | Unit, e2e per AC, coverage check, PII log scan on every commit | AC-3 (no card numbers in logs) is enforced forever, not checked once |
| Review | Every line reviewed | High tier, so the engineer reads the diff | Review effort matches risk |
| Release | Big bang | Flag on for 5% of customers, then ramp | A failure in AC-2 hits few customers |
The fixes compose. The prototype rules keep the branch safe to run. The decision log feeds the spec. The spec's numbered criteria make the coverage check possible. The coverage check makes it reasonable to trust agent-written tests. The risk tier from the log decides how much human review the loop needs. Remove any one link and the next one stops working.
How to tell whether it is working
Track a handful of numbers per feature. You do not need a dashboard for this on day one; a spreadsheet updated at each release is enough.
| Metric | Split payments | What a bad value tells you |
|---|---|---|
| Days from prototype to signed-off spec | 2 | Long: the prototype is wandering without a clear question |
| Days from spec to merge | 4 | Long: the engineer is discovering behaviour the prototype should have found |
| Decision records | 9 | Zero: decisions are living in chat |
| Acceptance criteria with an e2e test | 4 of 4 | Anything less: the loop is trusting untested behaviour |
| Production bugs a criterion should have caught (30 days) | 0 | Any: the spec is missing criteria, or QA is not reading tests |
| Share of prototype code kept after hardening | 55% | Near 100% on a high-risk feature: nothing was hardened |
Two of these deserve attention. If spec_to_merge_days is long, the engineer
is discovering behaviour that should have been found in the prototype, which
means stage 2 or 3 is thin. If prototype_lines_kept_pct is near 100 on a high
risk feature, the engineer probably did not harden anything.
What this does not fix
Bad product judgement. Prototyping faster means you can build the wrong thing faster. The loop verifies that the feature matches the spec. It says nothing about whether customers want it.
Deep technical work. A new payments ledger, a data migration, a latency problem in the search index: these do not start with a PM prototype. They start with an engineer, and the relay order (design the system, then build it) is still correct for them.
A weak test suite. If your existing product has no e2e coverage, the loop has nothing to stand on. Agents will happily write the first tests for you, but someone has to decide which ones matter.
Organisational trust. A PM running the app locally needs repo access, seed data, and permission to break things on a branch. Some companies are not ready for that, and no workflow document changes it.
The cost
Nothing in this shape is free.
- PMs and designers take on setup. Running a real app locally, managing branches, and reading agent output is a skill. Expect a few weeks where prototypes are slower than the old mockups.
- Seed data and fixtures become a product. Someone has to keep them realistic, or every prototype is tested against a fantasy.
- The engineer's job gets harder, not easier. Fewer engineers doing only the hardest part means more cognitive load per person. Hardening is less forgiving than building.
- The loop is infrastructure. Coverage checks, preview environments, risk tiers, and PII scans all need owning. On most teams, that owner does not exist yet.
The playbook
What to do, by situation.
| Situation | Who leads | Where the work happens | Human code review | What must exist first |
|---|---|---|---|---|
| Exploring a new problem | PM and designer | Miro or FigJam, then Claude Design | None | A problem statement |
| New user-facing feature, low risk (copy, layout, empty states) | PM and designer | Prototype on product, preview env | No; check the output in the product | Prototype rules, e2e per AC |
| New user-facing feature, high risk (money, data, permissions) | PM and designer, then engineer | Prototype on product, then hardening branch | Yes, engineer reads the full diff | Decision log with risk_tier: high, QA-signed spec |
| Internal tool or admin screen | Whoever needs it | Prototype on product | Spot check | Access controls reviewed once |
| Infrastructure, migrations, performance | Engineer | Editor with agents | Yes, plus a second engineer | A design doc, not a prototype |
| Bug in shipped feature | Engineer or PM | Failing e2e test first, then the fix | Matches the feature's tier | The AC the bug violated, or a new one |
| Prototype nobody wants to finish | PM | Decision log | None | A record saying why it was dropped |
And the adoption order, if you are starting from the relay:
- Write the prototype rules into
CLAUDE.mdorAGENTS.md, and protectproto/*branches. - Make seed data good enough that a PM can run the app locally in under ten minutes.
- Start a
decisions/folder and require a link to it in PR templates. - Adopt numbered acceptance criteria in specs, and have QA review them.
- Add the spec coverage check to CI.
- Add risk tiers, and only then relax code review for low-tier changes.
Do them in order. Step 6 without steps 4 and 5 is how teams end up shipping code nobody read and nobody tested.
This loop is not easy to set up
Look at the adoption list again. Every item is plumbing: agent instructions, seed data, decision records, spec parsing, coverage gates, risk-aware review. None of it is the product your team is trying to build, and all of it has to work before the three-person shape pays off. Most teams we talk to have the tools from the table above and are missing the loop underneath them.
That is the gap we are closing at CatalEx: making the verification loop that AI-native teams depend on something you turn on rather than something you build and babysit. We will share more on that soon. Stay tuned.
Written by CatalEx Engineering. We build the AI operating layer for AI-native companies: one platform to build, deploy, and run AI agents in production. More at catalex.co.