Skip to main content

If your team has adopted AI coding tools, you already know the uncomfortable math. Engineering teams that shipped a handful of pull requests a week are now shipping fifty, a hundred, sometimes more – much of it written by coding agents, reviewed by code review agents, and merged at a pace no manual QA process was designed for. One team we spoke with recently had just completed their first fifty PRs written entirely by agents and asked us the obvious question: what verifies all of this?

The traditional answer – write more test scripts – doesn't survive contact with that PR volume. Scripts take longer to author and maintain than agent-written code takes to produce, so the suite falls behind on day one and never catches up.

This guide walks through the alternative we built at QA.tech: connecting our QA agents to your repository so that every pull request triggers goal-based tests that derive themselves from what changed. No test code, no selectors, no maintenance queue. Here's how to set it up and what to expect at each step.

Step 1: Let the agent learn your application first

Before the PR integration does anything useful, the agent needs to know your product. When you connect QA.tech to your staging or sandbox environment, it runs an initial crawl: it explores the application the way a meticulous new hire would, following every path and action it can find, and assembles that into a knowledge graph of how your product is structured.

Two practical notes from hundreds of onboardings. First, if your application sits behind a login, set the login flow up as a prerequisite test case – the agent completes it first, then crawls everything behind it. Second, you control crawl depth. A shallow crawl (one or two steps deep into each user journey) is fast and usually enough to start; you can deepen it for the flows that matter most rather than exhaustively mapping everything up front.

This step is the whole reason the PR integration works. Because the agent knows every nook of the application, it can reason about what a given change might affect – which is exactly the judgment a good QA engineer applies when deciding what to check before a release.

The initial crawl maps your application's pages and flows into a knowledge graph, which every subsequent test builds on.

Get regular tech leadership wisdom for delivering better software and systems.

Step 2: Connect the repository

The GitHub integration takes minutes: authorize the app, pick the repositories, and choose when tests should fire – typically on pull request creation or update. GitLab merge requests follow the same pattern. From this point, testing is part of your pipeline rather than a stage that happens after it.

Step 3: Open a pull request and watch what happens

When a PR opens, the agent reads what was added and changed, cross-references that against its knowledge graph, and generates the test cases that change warrants – typically four or five for a routine PR, more for changes that touch shared flows. It then executes them immediately in a real browser, interacting with the updated build visually, exactly as a user would.

This is the part that surprises teams used to CI-triggered script suites: nobody authored those test cases. They were derived from the change itself, in the context of everything the agent already knows about your product. Ship a new filter on your dashboard, and the agent tests the filter and the dashboard flows it lives inside. Nobody had to remember to add coverage.

Test results appear directly in the pull request as a check, so reviewers see product-level verification alongside code review.

Step 4: Read the results where your engineers already work

Pass or fail lands in the PR itself. For every run, the agent produces an assessment report: the goal, the expected result, what actually happened, a video of the run, step-by-step screenshots, and the console and network logs underneath. A failure isn't a red X to reverse-engineer – it's a written explanation of what the agent expected to see and what it saw instead.

Because the agent tests through the interface, it reports failures at the product level – it will tell you the discount code no longer applies at checkout, with the evidence attached. That product-level context is exactly what your engineers – or your coding agents – need to locate and fix the cause. Teams running agentic development loops feed our assessment reports straight back to their coding agents as the failure context for the fix – and with the MCP integration, those coding agents can trigger test runs themselves while they work.

A failed run includes the agent's reasoning, a full video, and console and network logs – enough context to hand directly to an engineer or a coding agent.

Step 5: Layer regression on top of PR testing

PR-triggered tests verify the change. Regression plans verify everything else. In QA.tech you group test cases into plans – a smoke suite for every deployment, a full regression ahead of major releases – and run them on whatever cadence fits your release process.

Because tests execute in parallel, the wall-clock time is a fraction of what the same coverage costs manually. A recent full regression plan we ran covered a volume of testing that would occupy a manual tester for a day or two; it completed in about twenty minutes. That difference is what makes "run the full regression on every release" an actual policy instead of an aspiration.

Step 6: Ask the agent where your coverage is thin

Because the knowledge graph maps every flow the crawl discovered, the platform knows the difference between what your product contains and what your tests exercise. You can ask directly in the chat: "What are the highest-value flows we don't have coverage on right now?" The agent answers from its map of your product and can generate the missing tests on the spot.

This inverts the usual coverage dynamic. Instead of coverage being whatever accumulated over years of ticket-driven test writing, it becomes a question you can ask and act on in the same afternoon.

What this looks like after a month

The pattern we see across teams is consistent. Week one involves some back-and-forth – the agent asks questions, occasionally needs context about your product's quirks, and you correct its understanding. By week two the questions are rare. By week three they've mostly stopped, because the knowledge graph has absorbed how your product works. From there, testing runs at the pace of your pull requests, and the humans involved review results instead of maintaining scripts. If you want the full rollout sequence – from proof of concept through retiring your legacy suites – we've published it as a field guide, Adopting Agentic Testing: The Blueprint.

For teams that have made the jump to agent-written code, this closes the loop that was missing: code generation, code review, and verification all running at machine pace, with your engineers directing the system instead of feeding it.

FAQ

How many test cases does QA.tech generate per pull request?

A routine PR typically produces four to five test cases, derived from what changed. Larger changes that touch shared flows generate more.

Does PR testing require existing test cases?

No. The agent generates test cases from the change and its knowledge of your application. Teams starting from zero test coverage can begin at the PR level and build regression plans over time.

What happens when the UI changes in a pull request?

The agent works from goals and visual understanding rather than selectors, so it completes the test against the new interface and flags a failure only when the behavior itself is wrong.

Which CI/CD setups does QA.tech's PR integration support?

GitHub pull requests and GitLab merge requests are the primary integration points, with results reported back into the PR as a check.


See what QA agents generate for your next pull request – connect a repository at QA.tech and run the first test in minutes.

QA.Tech
By QA.Tech