Key idea
Agents are strongest on small, well-described, checkable tasks in code that looks like a lot of other code. They're weakest where the answer depends on something they can't see, and that's where they invent things.
Where they do well
- Common patterns. A form, a JSON endpoint, a login page, a test for a pure function. They've seen thousands of these.
- Mechanical changes. Renaming across files, updating call sites, turning a callback into
async/await. - Explaining code. "What does this file do?" is a cheap, low-risk question, and a good way into a project you didn't write.
- Anything with a fast check. If a test or a build tells it right from wrong, the loop from the last lesson does real work.
Where they struggle
Big, vague scope. "Make it production-ready" touches everything and has no finish line. The agent will change a lot of files and stop when it runs out of ideas, not when the app is ready.
Things outside the context. An agent sees only what's in its context: your request, the files it opened, and command output. A large project doesn't fit, so it works from a sample. It may miss the helper that already exists and write a second one, or break a caller in a file it never opened.
Long sessions. As a conversation grows, early instructions carry less weight, and older parts may be summarised or dropped. A rule you gave an hour ago can quietly stop applying. Short, focused sessions work better.
Your business rules. Whether a task due today counts as overdue is a decision, not a pattern. The agent will pick an answer; it may not be yours.
Invented APIs and packages
When a gap needs filling, a model can produce a function name, a config option or a package name that fits the pattern and doesn't exist. Sometimes the code just fails to run, and the agent notices. The dangerous case is a package name.
Attackers watch for names that models tend to invent and publish real packages under them, with malicious code inside. This is called slopsquatting. An agent that adds an invented name to your dependencies may install the attacker's package without an error in sight.
Before you accept a new dependency:
- Look the package up on its registry yourself. Check it's the project you expected, how old it is, and how widely it's used.
- Ask whether you need it at all. The starter app in this path has one dependency, Express, on purpose.
Module 6 comes back to this with the other security risks in generated code.
Check yourself