We redesigned aictrl's Backlog seven times in two days. AI agent testers and a simulated user built on Jev, TypeSafe's System One model, checked each version. No real users took part: novice agent testers stood in for first-time users, and the title says what those users would have done. This is the story of four moments where their evidence overturned what we believed, and of the places where it fell short.
We cut the Backlog landing screen from 37 controls to 16. It looked calmer. Then we ran the next test. Both novice agent testers approved a security review that nobody had asked about. Their task was to start a piece of work, and neither did.
The simplification caused the harm. Measurement caught it in one round. This post covers that round and three more like it, where a test result overturned what we believed about a screen.
Every number about our Backlog comes from our own measurement between 24 and 26 September 2026. Outside studies are cited where they appear, mostly in Part 5. The testers were AI agents and a simulated user, not people. We come back to what that means in the limits section.
Key findings
Before this work, Epics, Planner and Chat were three separate places in aictrl. You could not start a workflow from an issue. A run did not record which issue it was for. Our goal was to make the Backlog the control room: the one place where you start agent work and follow it.
AI agent testers drive the page through a Playwright harness. The harness exposes only the accessibility tree, which is what a screen reader gets. Haiku plays the novice, in a text variant and a vision variant. Sonnet is the control. A "hurried" persona answers first-glance questions.
We read success from the page state, never from the tester's own report.
The testers do tasks rather than judge screens. When Baymard Institute tested ChatGPT-4 on UX audits, 19.9% of its suggestions were accurate and 80.1% were false positives (Baymard, 2023). An off-the-shelf agent judging usability did worse than chance (AUC 0.436) until it was trained for the job (Gao et al., 2026). So our agents do the task, and the page state says whether they succeeded.
Why page state, not the tester's word
In the v4 round, a tester reported success after creating nothing. Its report said the task was done. The page said otherwise.
We wrote 22 tasks across 4 roles: engineering lead, product owner, reviewer and admin. Each task has a weight (frequency × criticality), a success check and a must-not check. The must-not check is how we see harm, such as an approval nobody asked for.
Each variant gets one score:
We accept a variant only if J improves and no criticality-3 task gets worse. A better average cannot hide a broken critical task.
An agent round takes about an hour. We wanted something faster for the first look, so we built a simulated user. Code does the perceiving. A scanner sees only what is above the fold: the top 14 controls by visual weight. A reader sees the whole page.
Jev makes one snap judgement per screen: which control next, am I done, can I answer. We sample many walks per task under four conditions: focused, laptop (1366×768), interrupted and paraphrased. Judgements are cached per screen, so a run over the task model takes minutes. One task costs about 100–200 Jev calls.
Each version answered a finding from the one before. Reviews shaped the early versions. Tests shaped the later ones.
The explorer has four views. Screen marks what changed. For v5, v5.1 and v6, Tester clicks traces where agent testers clicked, and Simulated first click shows where Jev expected a first-time user to click first. Compare puts a version and the one before it under a slider, or, for those three, shows the tested before and after side by side.

Landing screen: 10 controls, 98 words (a chat home, not a list).

Landing screen: 34 controls, 180 words.

Landing screen: 37 controls, 299 words.

Landing screen: 16 controls, 114 words.

Landing screen: 20 controls, 157 words.

Landing screen: 28 controls, 214 words.

Landing screen: 28 controls, 214 words.
main, the same way as in Figure 2. Watch where ready work sits: v5 drops it, v5.1 brings it back under approvals, v6 moves it to the top. Click data is available for v5, v5.1 and v6: where agent testers clicked, and where Jev’s simulated user predicted a first click, each on the screen that experiment used (1440×900 for v5 and v5.1; 1366×768 before and after moving Ready to start above approvals, for v6).main, measured the same way each time: 1440×900 viewport, default scenario. Controls use the left axis, words the right. Source: aictrl internal measurement, 24–26 Sep 2026.Density fell hard at v5. Then it climbed back as each fix added something a tester had needed. That climb is the subject of the limits section.
v5 applied four rules: one word per status, one verb, one place per job and one attention list. The attention list was Needs you. It opened on approvals and hid the work that was ready to start. Both Haiku novices approved a security review. Neither started the issue they were asked to start.
v5.1 put ready work back inside Needs you. Buttons now say what they launch. Blocked issues say what they wait on. Approvals ask for confirmation. All 3 novice testers started the right issue.
A cleaner screen is not a safer one. The first thing on screen is what a hurried first-time user would act on.
On a 1366×768 laptop, the first screen of v6 showed only approvals. The simulator flagged this before any agent run did.
We tested one change: move Ready to start above approvals. Over 24 walks per arm, success rose from 0.04 to 0.96 and harm fell from 0.67 to 0. The 90% confidence interval of the difference was +0.83 to +1.00.
Then we checked it with hurried agent testers on a laptop viewport. Before the change, 1 of 3 started the right issue. After it, 3 of 3 did. Wrong approvals fell from 2 to 0.
A tester triaging the backlog approved a gate nobody had asked for. We had two ideas.
Agent testers agreed with Idea B. Unasked approvals fell from 1 of 7 to 0 of 5. The one mistaken click was cancelled at the confirm.
The null result for Idea A mattered as much as the win for Idea B. It kept a change with no measured gain out of the design.
v6 was the first version built from the real product CSS and components. We ran the full role model against it. Success was 83%. After seven fixes it was 96%, and J fell from 4.27 to 2.47.
Several failures traced back to a stub or an icon-only control. The table shows six of the fixes and what they did. The seventh lets chat answer “who is working on X”.
| Task | Before | After | What changed |
|---|---|---|---|
| R3 Send a review back | 0/4 | 4/4 | A real Request-changes form replaced a stub |
| L7 Stop a run | 0/1 | 3/4 | Cancel run, with a confirm |
| P2 Split an issue | 1/4 | 4/4 | The proposal carries notes, and accepting it commits the split |
| L6 Pick a workflow | 2/4 | 4/4 | A visible radio list |
| P6 Change a status | 2/3 | 3/3 | A labelled menu instead of a "⋯" icon |
| L8 Who is working on what | 2/3 | 2/3 | An In progress group with owners; vision testers still miss it below the fold |
A simulated user built on Jev screened layout changes in minutes before agents confirmed them.before → after
We checked the simulator against 72 agent sessions.
Distracting, irrelevant text in the prompt changed nothing. The fold is what hides things, not a cut to the top few controls.
The rule we use
The simulator ranks variants and finds first-glance risks in minutes. Anything it finds, we confirm with agent testers before we accept it.
Our runs cover one product area over three days. Larger studies point the same way.
In one study, 49.0% of a GPT-4o simulated user's turns were polite, against 15.3% for humans (Zhou et al., 2026). The authors call simulators “excessively cooperative, stylistically uniform, and lack realistic frustration or ambiguity.” In another, agent success rates moved by up to 9 percentage points depending on which LLM played the user (Seshadri et al., 2026).
Kuric and colleagues compared GPT's predictions with 12 first-click tests and 3,431 real participants (Kuric et al., 2026). The predictions differed significantly in 53% of tasks. Personas and chain-of-thought “fail to create sensible fidelity improvements apart from inflating believability.”
Our Jev simulator makes the same kind of prediction. Against 72 agent sessions it caught the misclick and the fold problems, but missed multi-turn chat and interpretive answers. So we use it only to rank layout bets, then confirm with agent testers and, finally, people.
In UXCascade, UX professionals working from simulated agents found about as many issues as those working from a report written by people: 2.875 against 2.625 each, not a significant difference (Holter et al., 2026). The human report was stronger on discovery and interaction issues: 5 of 7 against the agents' 2 of 7. WebProber found 29 usability issues across 120 academic personal websites, many missed by traditional tools (Ye et al., 2025). At Userbrain, 5 of 5 synthetic users completed a task that 3 of 5 real users completed (Rössler, 2026). The simulation overestimated success.
Nielsen Norman Group sees synthetic users as preparation for research with people, not a replacement: “A synthetic user can be useful if researchers treat the output as a hypothesis to guide future research.” (Rosala and Moran, 2024). Raluca Budiu found that “the standard deviation in the synthetic data was consistently lower than the standard deviation in the human data.” (Budiu, 2025). Less spread means fewer edge cases. A MeasuringU review of experiments with synthetic users counted 9 encouraging findings and 14 discouraging ones (Lewis and Sauro, 2026). UX researchers rated UXAgent's simulated sessions 3.4 out of 5 for helpfulness and 3 for realism (Lu et al., 2025).
| Method | Good at | Misses | Evidence |
|---|---|---|---|
| One-shot AI audit (judges a screenshot) | Speed; it needs no tasks or harness | Most real problems; most of what it flags is wrong | Baymard (2023); Gao et al. (2026) |
| Simulated user predicting first clicks (our Jev simulator) | Ranking layout bets in minutes; first-glance and fold risks | Multi-turn chat and interpretive answers; real click patterns; it is too cooperative | Kuric et al. (2026); Zhou et al. (2026); this post |
| Agent testers doing tasks, success read from page state | Broken flows, stubs and unlabelled controls | Discovery issues; colour; they succeed where people may not | Holter et al. (2026); Ye et al. (2025); Rössler (2026); this post |
| People | The final answer: real spread, frustration and edge cases | Rounds are slower and cost more, so they run less often | Rosala and Moran (2024); Budiu (2025); Lewis and Sauro (2026) |
Agent testers are patient. They read text rather than layout. They never see colour. We treat their results as a check on clarity and flow. Real users give the final answer.
Our testers read the accessibility tree, which catches unnamed or mislabelled controls. That is no substitute for testing with people who use screen readers.
Erika Hall argues that talking to real people is not optional: “It is unethical, indefensible, and also unnecessary, to create a product or service or policy that affects other people, without having conversations with representatives of those populations.” (Erika Hall, LinkedIn)
Vitaly Friedman puts it as a trade: “I would always choose one day with a real customer instead of one hour with 1,000 synthetic users pretending to be humans.” (Friedman, 2025)
We agree on the final answer. Our agents check flows between rounds with people, not instead of them.
Since v5, controls went from 16 to 28 (+75%) and words from 114 to 214 (+88%). Each addition had a measured reason. The screen is still below v4's 37 controls and 299 words. The open question is whether 214 words is right. The next test is the real page, with people.
A 6-task baseline is 18 tester runs, mostly Haiku. A 20-task round is about 70. The simulator screens a hypothesis in minutes.
What would you test first on your own product?
No. Agent testers check clarity and flow between rounds with people, not instead of them: they are patient, read text rather than layout and never see colour. Nielsen Norman Group treats synthetic users as preparation for research with people, not a replacement, and a MeasuringU review of experiments with synthetic users counted 9 encouraging findings against 14 discouraging ones. People give the final answer.
Read success from the page state, never from the tester’s own report. Each of our 22 tasks has a success check and a must-not check, and the must-not check catches harm such as an approval nobody asked for. In one round, a tester reported success after creating nothing; only the page state showed the failure.
Accurate enough to rank layout bets, not to give the final answer. Checked against 72 agent sessions, our simulated user built on Jev caught the misclick and the fold problems but missed multi-turn chat and interpretive answers. Outside studies agree: in 12 first-click tests with 3,431 real participants, GPT’s predictions differed significantly in 53% of tasks (Kuric et al., 2026), and 49.0% of a GPT-4o simulated user’s turns were polite, against 15.3% for humans (Zhou et al., 2026).
A 6-task baseline is 18 agent-tester runs, mostly on Haiku, and a 20-task round is about 70. An agent round takes about an hour. The simulated user is the cheaper first look: it screens a hypothesis in minutes, at about 100–200 Jev calls per task.
Use the open-source ux-testing skill in our public skills repo on GitHub. It runs agent testers against your pages through a Playwright harness that exposes only the accessibility tree, and reads success from page state. Write tasks with a success check and a must-not check before you design, then confirm what the agents find with people.
Try it on your own product
The ux-testing skill in our public skills repo runs agent testers against your UI through a Playwright harness, reads success from page state and turns the results into a before-and-after report. Agent testing needs no API key. The optional simulated user needs a fast model API such as TypeSafe’s Jev (set TYPESAFE_API_KEY).
How aictrl.dev helps
Disclosure: aictrl builds workflow orchestration and skills for agent work, so we have a commercial interest in the practices described here. The ux-testing skill used in this post is published in our public skills repo at skills/ux-testing. Jev is TypeSafe's System One model (typesafe.ai), used under our own API key. To run design loops like this one as workflows, see how aictrl.dev can help.
main, 1440×900, default scenario).Published September 26, 2026 · Analysis by Bulat at aictrl.dev, co-authored with AI