Skip to main content

One PM, One Designer, One Engineer: How AI-Native Teams Build Now

· 22 min read
CatalEx Engineering
The team building CatalEx
CatalEx Engineering · Published September 12, 2026 · 09:00 UTC

It is Tuesday morning and a product manager is running the company's checkout page on their laptop. Not a mockup of it. The actual app, on localhost, on a branch called proto/split-payments, with seed data loaded. They ask Claude Code to let a customer split an order across two cards. Four minutes later they are clicking through it. The confirmation screen looks wrong when one card is a gift card with a low balance, so they call the designer over, and the two of them fix the layout in the running product while the agent keeps up.

By lunch they have found six edge cases nobody would have written into a requirements document, killed one idea (three-way splits), and written down why. No engineer has touched it yet. No Figma file exists. The engineer will be pulled in on Thursday, once the PM and designer are satisfied, and their job will not be "build this." It will be "make this safe to ship."

This is what a product team looks like in a company that builds with AI natively: one PM, one designer, one engineer, a set of agents, and roles that overlap so much they are close to interchangeable. The shape works. It is also much harder to set up than it looks, because the whole thing rests on a verification loop most teams do not have. That loop is the problem we are working on at CatalEx, and we will have more to share on it soon. Stay tuned.

This post is the playbook: how the new shape works, the tools teams use at each stage, where humans still read code and where they do not, what breaks, and what to do in each situation.

The old shape was a relay race​

For about fifteen years, product work moved through a relay. Each person produced an artifact whose only purpose was to be read by the next person.

  1. The PM wrote a PRD.
  2. The designer turned the PRD into mockups.
  3. The engineer turned the mockups into code, discovering along the way that the PRD skipped half the edge cases.
  4. QA tested the code against the PRD, which was by now out of date.

Every arrow in that chain is a lossy translation. The PRD cannot tell you that the confirmation screen overflows on a small phone. The mockup cannot tell you the refund API rejects partial amounts. You only learn those things by running the product, and in the relay, the first person who runs the product is the engineer, at the most expensive point to change your mind.

The team shape followed from the relay. You needed several engineers per PM because engineering was the slow, expensive stage where all the discovery happened.

The new shape: discovery moves to the product​

Coding agents changed which stage is expensive. Producing a working version of a feature is now cheap. Producing a correct, safe, maintainable version is still expensive. So AI-native teams split the work along that line.

In the relay, discovery happens in the red box, late and in code. In the new loop, discovery happens in the first box, early and in the product. The prototype is no longer a picture of the feature. It is the feature, running, unhardened.

That one change reshapes every role:

RoleIn the relayIn the AI-native team
PMWrites the PRD, answers questionsPrototypes on the product, finds edge cases by clicking, owns the decision log
DesignerDraws mockupsShapes look and feel in the running product, guards the design system
EngineerTranslates mockups into codeTakes a working prototype to production: failure modes, privacy, security, performance, observability
QAWrites test cases against the PRDReviews the spec and its acceptance criteria, owns what "done" means
AgentsAutocompleteWrite most of the code, tests, and docs at every stage

The roles converge because they all now work on the same artifact. The PM and the designer are both editing the running product. The engineer reads the same decision log the PM wrote. A designer who can run Claude Code can prototype without the PM, and a PM with taste can adjust spacing without the designer. What stays distinct is judgement, not tooling: the PM owns what the product should do, the designer owns how it should feel, and the engineer owns whether it holds up.

How the loop actually works​

There are five stages. Each has an owner, a place the work happens, and an artifact it leaves behind. The artifact is what makes the next stage possible without a meeting.

1. Explore on a canvas​

Before anyone touches the product, the team dumps the problem somewhere visual: user quotes, screenshots of competitors, rough flows, the metric they want to move. Most teams still do this in Miro or FigJam. The difference in 2026 is that the canvas has agents on it. Miro's Sidekicks and Flows can cluster sticky notes, draft user flows from interview transcripts, and pull in context from connected tools, so the brainstorm ends with a structured board instead of a photo of a wall.

Artifact: a board with the problem statement, the flows worth trying, and the questions nobody can answer yet.

2. Prototype on the real product​

This is the stage that changed the most. The PM and designer stopped designing in a separate tool and started prototyping in the product itself. They clone the repo, run it locally with seed data, create a branch, and ask a coding agent to build the flow.

Why the real product and not a blank canvas? Because the real product has the real constraints. It has the design system, the existing navigation, the actual data shapes, the loading states. A prototype built there shows you the edge cases a mockup hides: the empty cart, the expired session, the currency with no decimal places.

For visual exploration before touching code, some teams start in Claude Design, which reads a codebase's design system and generates interactive prototypes from conversation, then hand the chosen direction to Claude Code to build into the app. Others go straight to Claude Code, Codex, or Cursor.

The safety of this stage depends on rules the agent follows in prototype mode. Teams put them in the repository's agent instructions file (CLAUDE.md or AGENTS.md), so every PM and designer gets the same guardrails without having to remember them:

CLAUDE.md · Prototype mode

Applies to any branch named proto/*.

  • Never push to main. Never merge. Prototypes are reviewed, not shipped.
  • Put new behaviour behind the flag proto_<feature>, default off. So a prototype can be demoed on staging without affecting anyone.
  • Build only with components from src/ui. If a component is missing, stop and log it in decisions/<feature>.md instead of inventing one. Invented components are how a design system dies.
  • Use fixtures/ for data. Never call production APIs, never read real customer records. Prototypers are not trained to handle production data.
  • After every change, append one line to decisions/<feature>.md: what changed and why.

Artifact: a branch with a working, flagged, unhardened feature.

3. Research and log every decision​

Once the PM and designer are satisfied with how the feature behaves, they switch from building to documenting. This is where research happens: checking the prototype against interview notes, looking at analytics for how often the edge cases actually occur, and asking an agent to read the diff and list every behaviour the prototype introduced.

The core output is a decision log. Every choice made during prototyping gets a record, including the ideas that were killed. The killed ideas matter most, because six months later someone will propose three-way splits again, and the log is what saves the team a week.

Here is what one record looks like. It fits on an index card, and that is deliberate: a record that takes more than two minutes to write does not get written.

Decision record DEC-0142 · Split payments · September 8, 2026
FieldEntry
DecisionCap split payments at two methods.
Decided byPM and designer
WhyThree-way splits doubled the failure states on the confirmation screen, and none of the 14 customer interviews asked for more than two.
RejectedThree methods. The prototype branch is kept as evidence.
Open questionRefund to which method when one card has expired? Owner: engineer.
Risk tierHigh. It touches money, so an engineer must read the final diff.

Two fields do more work than they appear to. Open question hands unknowns to the engineer by name, so they become hardening work instead of production surprises. Risk tier is read later by the verification loop to decide how much human review the change gets.

From the prototype and the log, the agent drafts a spec: behaviour, constraints, and numbered acceptance criteria. The PM edits it. QA reviews it. The spec is now the source of truth, and the prototype is evidence for it.

specs/split-payments.md · Risk tier: high

Behaviour

  • A customer may pay with at most two methods (DEC-0142).
  • The first method is charged the amount the customer enters; the second is charged the remainder.

Acceptance criteria

  • AC-1: A gift card with a balance below the order total pre-fills as the first method with its full balance.
  • AC-2: If the second charge fails, the first charge is voided within 60s and the customer sees the "payment not completed" state.
  • AC-3: No card number, full or partial, appears in client logs.
  • AC-4: A refund on a split order returns funds to each method in proportion to its charge.

Artifact: a decision log, and a spec with acceptance criteria QA has signed off.

4. Harden: the engineer arrives​

Only now does the engineer take ownership. They are not translating anything. They read the spec, the decision log, and the prototype diff, and they ask a different set of questions from the ones the PM asked:

  • How does it fail? What happens when the payment provider times out between the two charges? Is AC-2 even achievable with this provider's API?
  • What does it leak? Which fields reach logs, analytics, or third-party scripts?
  • Who can abuse it? Can a user split an order so the second charge is for a cent, and does that bypass fraud checks?
  • Does it scale? What did the agent do that works on seed data and falls over on a customer with 4,000 orders?
  • Can we see it? Are there metrics and alerts for the new failure states?

Agents help here too, but in a narrower role. The engineer runs a security review agent over the diff, has an agent enumerate failure paths through the payment code, and asks for tests covering each one. The engineer decides which findings are real.

Often the engineer rewrites parts of the prototype. That is expected. The prototype's job was to be right about what; the engineer's job is to be right about how.

Artifact: a production-ready pull request, behind the same flag.

5. The verification loop​

The last stage is not a stage a person performs. It is a loop that runs on every change, and it is what lets this team avoid reading every line of agent-written code.

The piece most teams miss is the spec coverage check. Agents are very good at writing tests that pass. They are less reliable at writing tests that check the thing the spec cares about. So the loop enforces a link: every acceptance criterion in the spec must be named by at least one end-to-end test, and every test must name the criterion it covers. The check itself is a few lines of script: read the criterion ids out of the spec, look for each one in the test titles, and fail the build if any are missing.

A covered criterion is not the same as a well-tested one, so QA also reads what each test does. Here is what a good test for AC-2 walks through, in plain steps:

End-to-end test · AC-2: a failed second charge voids the first
  1. Open checkout with a $120 test order.
  2. Turn on split payment.
  3. Enter a valid test card as the first method.
  4. Enter a sandbox card that always declines as the second method.
  5. Press Pay.
  6. The screen shows "Payment not completed".
  7. The payment ledger shows no money held on the first card.

Step 7 is the one agent-written tests tend to skip. A screen can say "voided" while the money is still held on the customer's card. Test the system of record, not just the screen.

The gates run in the order that fails fastest, so a broken change is rejected in seconds rather than after a ten-minute browser run:

OrderGateFails when
1Unit testsAny function-level behaviour breaks
2Spec coverageAn acceptance criterion has no test naming it
3End-to-end testsThe product does not do what a criterion says
4Security and privacy scansA secret is committed, or card or personal data reaches a log
5Review gateThe decision log says high risk and no engineer has approved the diff

The last gate is where the risk tier from the decision log pays off. A low tier change needs a green loop and a look at the preview environment. A high tier change needs a green loop and an engineer's approval on the diff.

The tools, stage by stage​

The tool list moves fast, so read this as a snapshot of what teams are using in the second half of 2026, not an endorsement.

StageCommon toolsWhat they are used for
ExploreMiro (with Sidekicks and Flows), FigJamClustering research, drafting flows, keeping context on a shared board
Visual explorationClaude Design, FigmaInteractive prototypes grounded in the design system; Figma remains the home of the design system itself
Prototype on productClaude Code, Codex (CLI and app), CursorBuilding the feature in the running app on a local branch
Research and logClaude or ChatGPT with product analytics and tracker connectors (Linear, Jira, PostHog via MCP)Checking edge cases against real usage, drafting decision records and specs
HardenVS Code or Cursor with Claude Code or Codex, security review agentsReading diffs, running agents in parallel, inspecting their work
VerifyPlaywright, the unit test runner, GitHub Actions, preview environments, feature flagsThe loop that decides whether a change is acceptable

Does anyone still open VS Code?​

Yes, and the reason is instructive. The engineers still live in an editor, but they use it differently. They are not typing most of the code. They are running several agents at once, watching what each one does, reading diffs, and stepping in when an agent goes somewhere it should not. The editor is a supervision console.

The reason to keep one is the user. When a change is headed to customers and touches money, data, or permissions, someone accountable has to see the code. An editor with a good diff view is still the fastest way to do that.

And when does nobody look at the code?​

When the risk is low and the loop is trustworthy. A copy change, a new empty state, a tweak to an internal admin screen: for these, teams look at the output on the product itself, in a preview environment, and trust the loop for the rest. Reading the diff of a tooltip change adds cost and catches nothing the e2e test would not.

The mistake is not skipping code review. It is skipping code review without a loop that earns it. "We don't read the code" is a reasonable policy when the spec is reviewed, the acceptance criteria have tests, and the tier is low. It is a dangerous one when any of those is missing.

What goes wrong​

The shape fails in a small number of recognisable ways. Each has a check you can automate or make a habit.

Failure modeWhat it looks likeThe check that catches it
Prototype ships by accidentA proto/* branch gets merged because the demo looked finishedBranch protection rejects proto/* merges; hardening happens on a new branch
Decisions live in chatThe engineer asks "why two cards?" and nobody remembersPR template requires a decision log link; CI fails if a flagged feature has no decisions/ folder
Spec drifts from the prototypeThe spec says four behaviours, the prototype has sevenAn agent diffs prototype behaviour against the spec before QA review, and lists the extras
Verification theatreAgent-written tests pass but assert nothing the spec cares aboutSpec coverage check, plus QA reviews test titles against AC ids
Hardening skipped under pressure"It works on staging, ship it"The risk tier gates merge; high tier requires engineer approval
Design system erosionEach prototype adds a slightly different buttonPrototype rules forbid new components; lint fails on UI imports outside src/ui

Why the obvious fix fails​

The obvious move is to give every PM and designer Cursor or Claude Code and declare the team AI-native. Teams that do only this get a burst of excitement and then a mess.

Prototypes pile up as branches nobody can evaluate. Each one is a few thousand lines of agent-written code with no spec, no decision log, and no tests. The engineer is asked to "just clean it up," which is harder than building it from scratch, because they first have to reverse-engineer what the prototype was supposed to do. Engineering becomes the cleanup crew, the engineer burns out, and leadership concludes that PMs should not prototype.

The tools were never the hard part. The hard part is the connective tissue: the prototype rules, the decision log, the spec with numbered criteria, the coverage check, and the risk tiers. Without that, overlapping roles do not make a team faster. They make ownership unclear.

A worked example: split payments, before and after​

The same feature, built by the same company, a year apart.

StepRelay (2025)AI-native loop (2026)Why it matters
DefinePM writes an 8-page PRD over a weekPM and designer prototype on the product in two daysEdge cases are found by clicking, not guessed
DesignDesigner mocks 12 screens in FigmaDesigner fixes layouts in the running appThe gift card overflow is caught on day one
DecideDecisions scattered across comments9 records in decisions/split-payments/, 3 rejected optionsThree-way split is not re-litigated next quarter
SpecPRD goes stale during buildSpec with AC-1 to AC-4 extracted from prototype, QA signs offQA owns "done" before code is hardened
BuildTwo engineers, three sprintsOne engineer hardens in four daysEngineer time goes to failure, privacy, abuse
VerifyManual QA pass before releaseUnit, e2e per AC, coverage check, PII log scan on every commitAC-3 (no card numbers in logs) is enforced forever, not checked once
ReviewEvery line reviewedHigh tier, so the engineer reads the diffReview effort matches risk
ReleaseBig bangFlag on for 5% of customers, then rampA failure in AC-2 hits few customers

The fixes compose. The prototype rules keep the branch safe to run. The decision log feeds the spec. The spec's numbered criteria make the coverage check possible. The coverage check makes it reasonable to trust agent-written tests. The risk tier from the log decides how much human review the loop needs. Remove any one link and the next one stops working.

How to tell whether it is working​

Track a handful of numbers per feature. You do not need a dashboard for this on day one; a spreadsheet updated at each release is enough.

MetricSplit paymentsWhat a bad value tells you
Days from prototype to signed-off spec2Long: the prototype is wandering without a clear question
Days from spec to merge4Long: the engineer is discovering behaviour the prototype should have found
Decision records9Zero: decisions are living in chat
Acceptance criteria with an e2e test4 of 4Anything less: the loop is trusting untested behaviour
Production bugs a criterion should have caught (30 days)0Any: the spec is missing criteria, or QA is not reading tests
Share of prototype code kept after hardening55%Near 100% on a high-risk feature: nothing was hardened

Two of these deserve attention. If spec_to_merge_days is long, the engineer is discovering behaviour that should have been found in the prototype, which means stage 2 or 3 is thin. If prototype_lines_kept_pct is near 100 on a high risk feature, the engineer probably did not harden anything.

What this does not fix​

Bad product judgement. Prototyping faster means you can build the wrong thing faster. The loop verifies that the feature matches the spec. It says nothing about whether customers want it.

Deep technical work. A new payments ledger, a data migration, a latency problem in the search index: these do not start with a PM prototype. They start with an engineer, and the relay order (design the system, then build it) is still correct for them.

A weak test suite. If your existing product has no e2e coverage, the loop has nothing to stand on. Agents will happily write the first tests for you, but someone has to decide which ones matter.

Organisational trust. A PM running the app locally needs repo access, seed data, and permission to break things on a branch. Some companies are not ready for that, and no workflow document changes it.

The cost​

Nothing in this shape is free.

  • PMs and designers take on setup. Running a real app locally, managing branches, and reading agent output is a skill. Expect a few weeks where prototypes are slower than the old mockups.
  • Seed data and fixtures become a product. Someone has to keep them realistic, or every prototype is tested against a fantasy.
  • The engineer's job gets harder, not easier. Fewer engineers doing only the hardest part means more cognitive load per person. Hardening is less forgiving than building.
  • The loop is infrastructure. Coverage checks, preview environments, risk tiers, and PII scans all need owning. On most teams, that owner does not exist yet.

The playbook​

What to do, by situation.

SituationWho leadsWhere the work happensHuman code reviewWhat must exist first
Exploring a new problemPM and designerMiro or FigJam, then Claude DesignNoneA problem statement
New user-facing feature, low risk (copy, layout, empty states)PM and designerPrototype on product, preview envNo; check the output in the productPrototype rules, e2e per AC
New user-facing feature, high risk (money, data, permissions)PM and designer, then engineerPrototype on product, then hardening branchYes, engineer reads the full diffDecision log with risk_tier: high, QA-signed spec
Internal tool or admin screenWhoever needs itPrototype on productSpot checkAccess controls reviewed once
Infrastructure, migrations, performanceEngineerEditor with agentsYes, plus a second engineerA design doc, not a prototype
Bug in shipped featureEngineer or PMFailing e2e test first, then the fixMatches the feature's tierThe AC the bug violated, or a new one
Prototype nobody wants to finishPMDecision logNoneA record saying why it was dropped

And the adoption order, if you are starting from the relay:

  1. Write the prototype rules into CLAUDE.md or AGENTS.md, and protect proto/* branches.
  2. Make seed data good enough that a PM can run the app locally in under ten minutes.
  3. Start a decisions/ folder and require a link to it in PR templates.
  4. Adopt numbered acceptance criteria in specs, and have QA review them.
  5. Add the spec coverage check to CI.
  6. Add risk tiers, and only then relax code review for low-tier changes.

Do them in order. Step 6 without steps 4 and 5 is how teams end up shipping code nobody read and nobody tested.

This loop is not easy to set up​

Look at the adoption list again. Every item is plumbing: agent instructions, seed data, decision records, spec parsing, coverage gates, risk-aware review. None of it is the product your team is trying to build, and all of it has to work before the three-person shape pays off. Most teams we talk to have the tools from the table above and are missing the loop underneath them.

That is the gap we are closing at CatalEx: making the verification loop that AI-native teams depend on something you turn on rather than something you build and babysit. We will share more on that soon. Stay tuned.


Written by CatalEx Engineering. We build the AI operating layer for AI-native companies: one platform to build, deploy, and run AI agents in production. More at catalex.co.