Skip to content
You're viewing the v2 docs. Looking for v1?Go to v1 docs
Docs
Arcus
Popular questions
↑↓ to navigate↵ to selectesc to close
Book a demo

Test generation and quality

Generation starts in 3 ways:

  • A prompt on QI Home: you describe what to test and attach context. This creates an Adhoc session.
  • A sprint: Arcus reads the sprint’s stories from your connected project management tool and drafts test cases for each story’s acceptance criteria.
  • Mapped developer context: a pull request or coding session linked to a sprint or an Adhoc session triggers generation for it.

Once a sprint or session exists, the Arcus panel continues generating into it rather than starting a fourth way.

  1. Open QI Home.

  2. Enter what you want to test in the prompt box. Specific prompts produce specific test cases: name the feature, the roles, the variations, and the edge cases that matter.

  3. In the Add Context bar, attach the sources the generation should read. Attach them as described in Context Management.

  4. Click Send.

The new session appears on the Adhoc tab with the status “Generating test cases…”. When generation finishes, open the row and review what arrived. The prompt box carries the notice “Your data is secure and never used for AI training.”

The Sprints tab needs a project connected from Jira, Azure DevOps, Linear, or ClickUp.

  1. On QI Home, click the more options icon (⋮) and click Manage connected projects.

  2. Open the tool’s tab.

  3. Turn on the toggle beside the tool name.

  4. Select the project.

  5. Click Link project. Arcus reads the active sprints in that project.

Arcus then drafts test cases for each story’s acceptance criteria. Each sprint appears as a row on the Sprints tab, showing “Generating test cases…” while generation runs, and its page fills as it completes.

Sprint generation reads the stories plus what they link to (Confluence pages, Figma designs, and pull requests), provided the linked artifact is publicly accessible or reachable through a connected integration.

An Arcus panel docks to the left of every sprint and Adhoc session page. Use it when coverage is missing, a flow wasn’t generated, or you want the plan adjusted, without leaving the page.

  1. Open the sprint or Adhoc session from QI Home.

  2. In the Arcus panel, describe what you want in “Ask Arcus about this cycle…”. Click the plus icon (+) to attach more context.

  3. Click Generate with AI.

The panel shows the generation as it runs: an “Orchestrating a generation plan” log with steps such as “Step groups searched”, “Existing tests searched”, “Scenarios validated”, “Plan updated”, and “Summary generated”. Arcus searches your existing library and step groups first, so coverage you already have isn’t duplicated.

When it finishes, the panel offers 3 follow-ups: Satisfied — finish, Drill deeper, or Additional work. Pick one, or reply in your own words.

However generation started, the new test cases arrive in Pending status and the Test Case Review card updates its count. Nothing enters your library until you accept it.

Accepting moves coverage immediately; coverage is the only metric that rises before any run. Whatever stays pending is counted as a gap.

  1. On the sprint or Adhoc session page, click Review test cases on the Test Case Review card. The card also lists the highest-value cases under Start with these; click one to open it directly.

The review page carries 2 cards: the current Coverage, and Pending review with the count still waiting, plus how many are approved and how many were already in your library.

Four controls above the list narrow what you review:

  • By: group the list by Story or by Modules impacted. Each group header shows the story or module, its coverage, its open gap count, its test count, and an Accept all button.
  • Gap type: filter gaps by where they came from: Requirement (generated from stories and attached context), Agentic Learning (found while exploring the live app), or CLI (pushed from a developer’s terminal).
  • Filters: narrow by Status (All, Pending, Accepted, Rejected), Story, Module, Priority, Origin, and Type.
  • Search tests: find a test case by name.

The Simple and Detailed toggles switch how much each row shows; the detailed view adds the test type.

  1. Click a test case name. The detail view opens with its description, preconditions, and steps with expected results.

  2. Move through the list with the arrows at the top; the counter shows where you are, such as “1 of 18”.

A step showing Steps added reuses a step group. Expand it to see the sub-steps, or open the step group itself from the link. Steps can be parameterized with data set variables such as $from_account.

To change anything, click the more options icon (⋮) and click Edit. An accepted test case stays editable from your library.

  • One test case: open it and click Save to Library, or Reject if it isn’t accurate.
  • A group: click Accept all on a story or module header.

Accepting is what moves the numbers:

  • Accepted test cases enter your library, count toward coverage, and can be automated.
  • Pending test cases are open gaps. They hold coverage down until you decide.
  • Rejected test cases are excluded from coverage entirely, and stay listed in the review.

To automate an accepted test case without leaving the review, click Explore and Automate in its detail view. See Explore and Automate.

A test plan groups test cases into runs. Until a plan runs, Pass Rate, Confidence, and Readiness stay empty.

The 4 tiers widen in scope:

  • Smoke: critical path (P0 and high-risk cases).
  • Feature: new functionality in this sprint.
  • Regression: modules impacted by the sprint.
  • Deep Regression: end-to-end, across every app type.

Every smoke case is contained in the feature tier, so running the wider tier validates everything the narrower one would have.

  1. On the sprint or Adhoc session page, click Generate test plan on the Plans & Execution card.

  2. Under Select a plan, select 1 or more tiers.

  3. Click Generate with AI.

A plan can include both accepted and under-review test cases. The selection screen shows the split, such as “1 accepted · 17 still under review”.

Each generated plan gets its own tab on the Test Plans & Execution page and waits in Pending Review until you act on it.

  1. Open the plan’s tier tab.

  2. Read the plan card: an Arcus suggested badge, the plan’s description, its execution count, its run count, and the bar splitting results into Passed, Failed, Not run, and gaps to review.

  3. Click Accept Plan to accept it and start its runs, or Reject.

The status line says what acceptance triggers, such as “Accept to start 10 runs”. An accepted plan shows Plan Accepted.

A plan’s under-review test cases surface as its gaps, and the card shows “39 gaps to review before the next run”. Review them first, so the next run’s results cover them.

Each plan splits into 1 test run per module, listed under Test Runs. A run row shows its machine, its test count, and its open gap count, with result counts on the right. Expand a run to see its test cases, each with its module, type, automation state, and latest result.

You refine a plan by asking for the change, not through a form.

  1. Click Modify plan, or + Create new plan. The Arcus panel opens.

  2. Describe the change (generate another tier, adjust coverage, reassign a run’s machine) in “Ask Arcus to plan, or refine an existing plan…”, and send it.

The panel reports what changed: the tier, the case count, the run count, the estimated duration, which signal sources it scored candidates on, and whether every smoke case still sits inside the feature tier. The updated plan appears on its tier tab, pending your acceptance.

Link a plan to a Testsigma automated test plan, so automated run results flow into the metrics.

  1. Open the plan’s tier tab.

  2. Click Link plan beside Automated Test Plan on the plan card. The Link Testsigma Automated Test Plan dialog opens.

  3. Select the Project, Application, Version, and Test Plan.

  4. Click Link.

The card then shows the linked plan beside Automated Test Plan, such as “TS-67 NA Sprint 1 | Smoke Testing”. Linking is per tier: link the Smoke plan to a Testsigma smoke plan, the Feature plan to a feature plan, and so on.

The QI (Quality Intelligence) Metrics are the 4 values in the sprint or Adhoc session header (Coverage, Pass Rate, Confidence, and Readiness) and the release gate verdict built on them. Thresholds are configurable per project under Settings > QI Metrics.

The metrics build on each other in order. Coverage reports how much of the generated surface you accepted. Pass Rate reports how much of what ran passed. Confidence reports how trustworthy those results are. Readiness rolls the 3 into a single score, and the release gate turns the checks into the verdict: Ready, Conditional Ready, or Not Ready.

The share of the generated surface that has been signed off. It rises as you review test cases, before any run.

Coverage = Accepted ÷ (Accepted + Pending) × 100

Pending test cases count against it. Rejected test cases are excluded entirely. Coverage is live from the moment test cases are generated and updates in real time as you accept. Default target: 85% or above.

The share of executed units that passed, where a unit is 1 test case on 1 machine. Manual and automated results are tracked separately, and each unit counts once, on its latest terminal result. A terminal result is one that finished: passed, failed, blocked, or skipped. Not run and in progress are excluded.

Pass Rate = Passed Unique (Test Case × Machine) Units ÷ Total Unique (Test Case × Machine) Units with Terminal Status × 100

Pass Rate shows Awaiting execution until a test plan runs. Default target: 95% or above.

The priority-weighted share of the planned scope that ran trustworthily: executed, and not flaky. A flaky test case is one that has both passed and failed across runs without its steps changing. Because weights follow priority, flaky P0 test cases cost the most.

Confidence = Trustworthy Weight ÷ Total Weight × 100

Confidence shows Awaiting execution until runs report. It lands in 1 of 3 bands. You set where High and Medium start; Low is whatever falls below Medium.

BandDefaultMeaning
High80% of score or aboveTrustworthy signals
Medium50% to 79% of scoreSome issues
LowBelow 50%, auto-derivedUnreliable signals

Confidence is scoped to what was planned. If you generate a wide tier and run only a narrow one, the unrun scope stays in the denominator, and the score reports the validation you committed to but did not perform.

Pass Rate, Coverage, and Confidence rolled into a single 0 to 100 shippability score. It stays empty until a run reports. The header card is labeled Readiness; Settings shows the same metric as Release Readiness.

Release Readiness = (Pass Rate × 0.40) + (Coverage × 0.25) + (Confidence × 0.35)

Pass Rate holds the largest weight because an active failure is a more direct signal than a known gap. Confidence modifies how far those results can be trusted.

The verdict on the sprint or Adhoc session (the status chip beside its name, and the Readiness column on QI Home) comes from 5 checks. Each check lands in Ready, Conditional, or Not Ready, and the worst result across all checks is the verdict. One Not Ready check overrides everything else.

CheckReadyConditionalNot Ready
Pass rate (tests passing)≥ 95%≥ 90%Below 90%
Coverage (tests approved)≥ 85%≥ 70%Below 70%
Confidence (signals trustworthiness)HighMediumLow
Critical flaky (P0 tests unstable)012 or more
Critical failures (P0 tests failing)012 or more

The page states the failing reason in plain terms: “Not ready to release. Coverage is 24%, below the 85% bar.” A sprint or Adhoc session with no run data yet shows Awaiting Run instead of a verdict.

ValueCases generated, none acceptedAccepted, no runsRuns executed
Coverage0%, liveUpdates as you acceptUpdates as you accept
Pass RateAwaiting executionAwaiting executionPercentage
ConfidenceAwaiting executionAwaiting executionPercentage with band
ReadinessAwaiting executionAwaiting execution0 to 100 score
VerdictAwaiting RunAwaiting RunReady, Conditional Ready, or Not Ready

Editing QI Metrics requires admin access.

  1. Go to Settings > QI Metrics.

  2. Review each metric’s formula.

  3. Set the Ready and Conditional thresholds for the 5 release gate checks. The Not Ready boundary is derived from them.

  4. Under Confidence, adjust where the High and Medium bands start.

Why is Coverage 0% when test cases were generated?

Section titled “Why is Coverage 0% when test cases were generated?”

Coverage counts accepted test cases, not generated ones. Generated test cases start in Pending and hold coverage down until you accept them. Coverage updates in real time as you do.

Pass Rate needs at least 1 executed test plan. Until a run reports, it shows Awaiting execution; accepting test cases moves Coverage only. Generate and accept a plan to start runs.

Why is the verdict Not Ready when I accepted test cases?

Section titled “Why is the verdict Not Ready when I accepted test cases?”

Because the worst of the 5 checks wins, and most checks need run data. Accepting test cases raises Coverage, but Pass Rate, Confidence, and the critical-failure checks move only after a plan runs. The line under the verdict names the failing check.

Was this page helpful?