Key idea
A coding agent is a language model working in a loop: it plans, edits files, runs a command, observes the output, and goes round again. It stops when it believes the job is done, which isn't the same as the job being right.
Chat assistant or agent
A chat assistant writes code into a conversation and you copy it out. An agent works inside your project: it opens files, changes them and runs commands in a terminal.
Claude Code, Cursor, Copilot's agent mode, Codex and Windsurf all work this way. App builders such as Lovable, Bolt, v0 and Replit run a similar loop behind a friendlier screen. Everything in this path applies to all of them.
The loop
The agent's loop, and your gates
Decides what to change
You approve the planWrites it into your files
You read the diffRound again until nothing it runs fails
Reads what came back
You decide when it’s doneTests, a build, the app
It plans from your request and the files it thinks matter, and it runs whatever it can: the tests, a build, the app, a linter. An error sends it round again; clean output ends the loop.
Observing is where the power comes from. An agent that can run your tests fixes its own typos and wrong imports without you.
It's also the limit. The agent only sees what it runs. If no test checks how overdue dates work, a wrong rule for overdue dates passes quietly, and the agent reports success.
Why agents are confidently wrong
The model predicts code that looks likely, based on patterns from a huge amount of code. Likely and correct overlap most of the time, which is why agents are useful. When they don't overlap, the wrong answer reads exactly as fluently as a right one.
You'll meet three shapes of this again and again:
- The wrong problem. It solves what it inferred from your request, not what you meant.
- An unchecked claim. "All tests pass" when it ran some of them, or none that cover the change.
- A plausible gap-filler. A function, setting or package that sounds right and doesn't exist.
So read "Done" in an agent's summary as "the loop stopped". Treat it like a colleague saying "should work": a claim you check before it ships.
Where you sit in the loop
The diagram's three gates are where you check each loop: you approve the plan, you read the diff, you decide when it's done. This path builds the habits behind them: a written brief to approve the plan against (module 3), reading every diff (module 4), running the tests yourself before you call it done (module 5), and approving anything that touches your live app (modules 8 and 9). The agent does the typing; you decide what ships.
What the agent can actually see
An agent works from its context: your request, the files it opened, the output of commands it ran, and any instruction files in the project. Anything else, such as last week's session, a decision you made in a meeting, or a file it never opened, doesn't exist for it.
That's why the same agent can be sharp on one task and lost on the next. Module 7 covers giving it the context it needs on purpose.
Check yourself