Good Gherkin Examples vs Bad Scenarios (with Fixes)
Good Gherkin Examples

Get Started

with $0/mo FREE Test Plan Builder or a 14-day FREE TRIAL of Test Manager

The gap between good Gherkin examples and bad ones comes down to one thing: behavior over button-clicks.

  • Good Gherkin describes what the system does. Bad Gherkin scripts how a user clicks through the UI and breaks the moment that UI changes.
  • Most broken scenarios share four flaws: vagueness, imperative steps, multiple behaviors crammed into one scenario, and copy-paste repetition.
  • Every bad example here comes with a paste-ready fix, plus a quick-reference table you can save as a cheat sheet.
  • Clean Gherkin matters more in the AI era because generated scenarios are often "almost right" and still need a human eye.

Treat each scenario as a contract for one behavior, and your feature files turn into living documentation instead of brittle scripts.


Writing good Gherkin examples is about describing behavior so clearly that a developer, a tester, and a product owner read the same scenario and picture the same thing. That skill is essential now that 84% of developers use or plan to use AI tools to help write code and specs. AI-generated scenarios tend to look correct while hiding subtle gaps.

This guide puts good and bad Gherkin side by side, explains exactly why each weak scenario fails, and hands you a fix you can paste in today. Pairing these habits with a modern test management platform built for writing and running Gherkin tests turns a tidy feature file into a quality workflow your whole team can rely on.

What Makes Good Gherkin Examples Good (and Bad Ones Bad)?

A scenario is "good" when it survives change. Business rules stay stable for years, while the UI gets redesigned every quarter. So good Gherkin examples describe rules and outcomes, not screens and selectors. Bad ones do the opposite, which is why they rot fast and fail on every refactor.

The One Principle Behind Every Good Scenario: Behavior, Not Buttons

The single rule that fixes most weak scenarios is to write declarative steps instead of imperative ones. A declarative step says "the customer requests a refund." An imperative step says "the user clicks the button with id refund-btn." The official Cucumber guidance puts it simply: your scenarios should describe the intended behavior of the system, not the implementation. Ask yourself one question for every step: Would this wording need to change if the implementation changed? If yes, rewrite it.

Given, When, Then: State, Action, Outcome

Each keyword has a job, and mixing them up is where readability dies. Given sets up a state that already exists before the behavior runs. When triggers the single action you're testing. Then asserts an observable outcome, and nothing more. A Given that performs an action is a When in disguise, and a Then that changes state is hiding the real assertion. Keep one When per scenario, and you've avoided half the mistakes in this article.

3 Marks of good gherkin scenario

Good Gherkin Examples vs Bad: 5 Side-by-Side Fixes

Theory only goes so far, so here are five common failures with the exact fix for each. Read the bad block, note why it fails, then steal the rewrite. These examples cover the patterns that show up in almost every messy feature file.

Fix 1: Vague Steps Become Specific Behavior

Vague scenarios are the most common bad pattern. They use words like "log in" and "logged in" that mean nothing concrete.

# Bad

Scenario: Login

    Given I am on the site

    When I log in

    Then I am logged in

Why it fails: nobody can tell what to build or what to assert. "Logged in" could mean a redirect, a welcome banner, or a session cookie. Here's the fix.

# Good

Scenario: Registered user signs in with valid credentials

    Given a registered user is on the login page

    When the user signs in with valid credentials

    Then the user lands on their account dashboard

Fix 2: Imperative UI Steps Become Declarative Behavior

This one glues every step to the current interface. It reads like a click-by-click script.

# Bad

Scenario: Free subscriber views an article

    Given I open the browser

    And I navigate to "/login"

    When I type "user@test.com" in the "#email" field

    And I type "Pass123" in the "#password" field

    And I click the button with id "submit"

    Then I see "Free Article" in the list

Why it fails: Change a field id or a route, and the spec breaks, even though the behavior never changed. The declarative version hides those details in the step definitions where they belong.

# Good

Scenario: Free subscribers see only free articles

    Given a user with a free subscription is signed in

    When the user opens the article library

    Then only free articles are visible

Gherkin best practices

Fix 3: Many Behaviors Become One Behavior Per Scenario

A scenario should test exactly one thing. This example stacks several actions and outcomes together.

# Bad

Scenario: Checkout

    Given the cart has items

    When the user checks out

    Then payment succeeds

    When the user checks out again with a bad card

    Then an error appears

Why it fails: Two Whens means two behaviors, so when the test goes red, you can't tell which one broke. Split it into independent scenarios.

# Good

Scenario: Successful checkout confirms the order

    Given the cart has one item

    When the customer completes checkout with a valid card

    Then the order is confirmed

Scenario: Checkout with a declined card shows an error

    Given the cart has one item

    When the customer completes checkout with a declined card

    Then a payment error is shown

Fix 4: Repetitive Scenarios Become a Scenario Outline

When you find yourself copy-pasting a scenario and changing one value, that's the signal to parameterize. A Scenario Outline with an Examples table runs the same behavior across many inputs.

# Good

Scenario Outline: Cart discount applies above the threshold

    Given a cart subtotal of <subtotal>

    When the customer views the order summary

    Then the grand total is <total>

    Examples:

| subtotal  | total      |

| 50.00     | 50.00    |

| 100.00   | 100.00  |

| 101.00   | 90.90    |

This pattern expands coverage without bloating your feature file, and you update the structure in one place when rules change.

Fix 5: Repeated Setup Moves into Background

If every scenario in a file starts with the same two Given steps, pull them into a Background block so they run once before each scenario.

# Good

Background:

    Given a signed-in customer

    And an empty cart

Scenario: Adding an item updates the cart count

    When the customer adds a product to the cart

    Then the cart shows one item

One caution: use Background only for static preconditions like identity or configuration. Don't put data-changing actions there because shared mutable state across scenarios makes test suites flaky.

What Are the Most Common Bad Gherkin Examples to Avoid?

Beyond the five fixes above, a handful of anti-patterns show up repeatedly in code reviews. Spotting these quickly is the fastest way to clean up an existing suite. Here are the bad Gherkin examples most worth hunting down and rewriting:

  1. Procedural scripts wearing BDD clothing. Putting Given/When/Then in front of a traditional step-by-step test does not make it behavior-driven. If the steps describe keystrokes, it's still procedure-driven.
  2. Mixed first- and third-person voice. Switching between "I" and "the user" inside one scenario creates confusion about who is acting. Pick third person and stay consistent.
  3. CSS selectors and endpoints leaking into steps. The moment a Given mentions #submit-btn or /api/v2/cart, the scenario is testing implementation, not behavior.
  4. Meaningless titles. "Test 1" or "Verify login" tells a teammate nothing during test triage. A good title names one distinct behavior in a single line.
  5. Then steps that change state. A Then should assert, never act. If your Then is creating data, you've buried the real check.
  6. Giant feature files with no structure. Thirty scenarios in one file with no tags or grouping is a maintenance trap. Group by business capability and tag for selective runs.
  7. Scenarios written after the code ships. Skipping the discovery conversation defeats the whole point. Peer-reviewed research on BDD found that writing scenarios collaboratively before implementation reduces misunderstanding and cross-role misalignment early in the project.

For a more exhaustive rundown, this companion piece on Gherkin anti-patterns to avoid digs into each failure mode in detail.

Here's the whole fix-it list in one place. Save this table as your cheat sheet and run it against any feature file before you commit.

Bad Gherkin patternWhat goes wrongThe fix
Vague steps ("I log in")No one can build or assert itName the actor, action, and visible outcome
Imperative UI stepsBreaks on every UI changeWrite declarative, behavior-level steps
Many behaviors per scenarioCan't tell what failedOne When, one behavior, split the rest
Copy-paste scenariosDrift and duplicationScenario Outline with an Examples table
Repeated setup linesNoise and fragilityMove static preconditions to Background
Meaningless titlesSlow triage and reviewOne clear sentence per behavior

How Do You Keep Gherkin Test Cases Clean as Your Suite Grows?

Writing one good scenario is easy. Keeping a thousand of them readable, runnable, and trustworthy is the real challenge, and it's where many teams give up on BDD. The Gherkin best practices in this guide build on the fundamentals covered in this complete guide to writing Gherkin tests. The next step is connecting those scenarios to execution so they don't drift into documentation theater.

Gherkin test cases

Connect Scenarios to Execution and Traceability

Feature files that live in a repo and never run quickly become stale. To stay useful, your Gherkin test cases need to map to actual test runs, link back to requirements, and show pass or fail history. That means importing your feature files into a system that can link scenarios to your GitHub pull requests and trace them to issues, so a failing scenario points straight at the change that caused it. When your specs drive real test management instead of sitting in a folder, they finally earn the "living documentation" label everyone promises.

From feature file to living documentation

Where AI Fits into Writing Gherkin Test Cases

AI is now writing a lot of Gherkin, and that raises the stakes on review. The developer survey found that 66% of respondents struggle with AI solutions that are almost right, but not quite. A generated scenario will mix imperative steps, double up behaviors, or invent a vague Then, all while looking polished. The smart move is to use AI to draft fast, then apply the fixes in this guide. Modern platforms are shifting toward an active intelligence layer, with QA agents that help generate behavior-focused scenarios with AI from a user story, output them in clean Gherkin, and flag the weak spots before they reach your suite. You write less boilerplate, and you still own the quality bar.

Good Gherkin Examples: Frequently Asked Questions

What's the difference between good and bad Gherkin examples?

Good Gherkin examples describe behavior and outcomes in domain language that survives UI changes. Bad ones describe clicks, fields, and implementation details, so they break constantly and read like scripts. If a step would need rewording when the interface changes, it's leaning toward bad.

How long should a Gherkin scenario be?

Short. Most strong scenarios fit in three to seven steps with a single When. If yours runs longer, you're probably testing more than one behavior and should split it. Background blocks and Scenario Outlines help you cut length without losing coverage.

Should Gherkin steps use first or third person?

Third person, consistently. Modern apps are multi-user, so "I" gets ambiguous about which actor is doing what. Phrases like "the customer" or "an admin" keep the role explicit and the scenario easier to reuse.

When should I use a Scenario Outline instead of separate scenarios?

Use a Scenario Outline when the same behavior runs across multiple data inputs, like discount tiers or validation rules. Use separate scenarios when the behaviors genuinely differ, such as a successful path versus an error path. The rule of thumb: same steps, different data means an outline.

Can AI write good Gherkin test cases for me?

AI can quickly draft solid Gherkin, and it's great for first passes and edge-case ideas. It also produces scenarios that look right but hide vague or imperative steps, so treat the output as a draft. Pair an AI test case generator with the Gherkin best practices and fixes in this guide, and review every scenario before it lands in your suite.

Ship Cleaner Scenarios Starting With Your Next Feature File

Clean Gherkin is a habit, and the fastest way to build it is to write scenarios your whole team can see, run, and trace in one place. That's where an AI-powered QA platform earns its keep. TestQuality pairs structured test management with QA Agents and TestStory.ai, so you can generate behavior-focused scenarios with AI, drop them straight into test plans, and link every scenario to your GitHub pull requests and Jira issues. It works the same whether your code is human-written or AI-generated, surfacing the quality signals you need to ship with confidence. Start your free TestQuality trial and turn your next feature file into living, executable documentation. 

Newest Articles

Diagram showing requirements traceability flow from Levr epic to TestStory.ai generated test cases to TestQuality execution
From Levr Epic to Executable Test Case: The Full Agentic SDLC
Requirements traceability is the practice of linking high-level business goals and acceptance criteria directly to the code commits, test suites, and execution histories that verify them. In modern software engineering, achieving requirements traceability confirms that autonomous coding agents do not introduce silent bugs or build misaligned features; this forms the core of a disciplined agentic… Continue reading From Levr Epic to Executable Test Case: The Full Agentic SDLC
Diagram showing a Cucumber feature file, step definitions with a custom World, and Playwright execution converging into a JUnit XML report uploaded via the TestQuality CLI
Agentic Testing and QA with Playwright and Cucumber
Playwright and Cucumber is a BDD testing combination that pairs Gherkin feature files with Playwright's browser automation engine, letting teams write behavior in plain language while executing it through deterministic, code-driven checks. Cucumber parses Feature and Scenario statements written with Given, When, and Then keywords, step definitions bind those statements to TypeScript or JavaScript functions,… Continue reading Agentic Testing and QA with Playwright and Cucumber
Circular diagram showing a Gauntlet Loop AI workflow — lead agent, builder agents, and judge agent connected in a loop, with output syncing to TestQuality
Gauntlet Loop: How AI Agents Build, Judge, and Fix Work
A Gauntlet Loop is an agentic AI workflow in which a lead agent breaks a broad goal into small, independently judgeable pieces, assigns them to specialist builder agents, and routes every result through a separate judge agent that compares the work against a quality bar. The pattern was popularized by Matt Shumer's July 2026 "Claude… Continue reading Gauntlet Loop: How AI Agents Build, Judge, and Fix Work

© 2026 Bitmodern Inc. All Rights Reserved.