Learning paths / Vibe coding to production / Tests that prove something

Let the agent write tests, then check them

Reading · 6 min · Module 5, lesson 2 of 311 min left in this module

Module 5 · Tests that prove somethingLesson 2 of 3

Goal: Spot agent-written tests that can't fail, and check any test by breaking the code it guards and running the suite yourself.

Key idea

Agents write tests quickly, and you should let them. But the agent writing the tests has the same blind spots as the agent writing the code. Check a test the only way that counts: break the code it guards and see it fail.

Tests that test nothing

These pass every time, whatever the code does. You'll see all of them in agent diffs.

  • The expected value came from the code. The agent runs isOverdue on a task due today, sees true, and writes a test expecting true. The test now protects the bug.
  • The test was changed to pass. A test fails, and the fix is a new expected value, with a summary like "updated test to match new behaviour". Sometimes that's right. It always needs a reason you agree with.
  • Nothing is asserted. The test calls the function and checks nothing, or wraps everything in try and ignores the error.
  • It never runs. test.skip, test.todo, or a file outside the folder the test runner reads.
  • It only proves the author's idea. The export tests from lesson 4.4.2 check commas and quotes, pass, and miss the formula cell entirely.

Here's the first kind in the starter, before your fix. It reads like a real test:

test("a task due today is overdue", () => {
  const now = new Date("2026-10-01T15:30:00Z");
  assert.equal(isOverdue({ due: "2026-10-01", done: false }, now), true);
});

It passes, and it's wrong. Only you know the rule; the name gives it away once you read it as a rule.

Four checks, in order

  1. Read the test names as rules. Each name should match an acceptance criterion from your brief. A name you disagree with is a bug, in the test or in the code.
  2. Break the code on purpose. Undo the fix by hand, say < back to <=, and run npm test. The new test must fail. Put the fix back.
  3. Run the suite yourself. Read the summary lines, not the agent's account of them: pass, fail, and also skipped and todo, which should be 0 unless you know why.
  4. Read every changed expectation. In the diff, any existing test whose expected value changed gets the same scrutiny as production code.

Step 2 takes a minute and catches most tests that can't fail.

Mutation testing: step 2, automated

Tools called mutation testers do step 2 for you, at scale: they make many small changes to your code, such as flipping < and <=, and report every change your tests didn't catch. They're slow on big projects, but on a small app they're a sharp check of whether an agent's test suite actually guards anything. Try one when a project's tests matter more than their count.

Ask for tests that can fail

Put the check in the request, so the agent does step 2 first and you repeat it:

Write tests for each acceptance criterion in the brief.
For every test, change the code so the test should fail, run it,
show me the failure, then restore the code.
Don't change any existing test without asking me first.

Then run npm test yourself, in your own terminal, before you merge.

Check yourself

The agent's summary says: 'One test was failing, so I updated its expected value. All tests pass now.' What do you do?
How do you check that a new test really guards your fix?