Coding-agents

How to use coding-agents effectively

Context windows, CLAUDE.md, plan mode, model selection — everything you need to work with AI coding agents like a senior engineer, not a beginner.

7 min read

The Mental Model Nobody Gives You

Most people treat AI coding agents like a search engine. Type a vague question. Get an answer. Move on.

That's why most people get mediocre results.

An AI coding agent — Claude Code, Codex, Cursor — is not a search engine. It's a language model running in a loop with tools. It reads your instruction, picks an action (read a file, run a command, edit code), observes the result, and repeats until it thinks it's done.

Think of it as a fast, capable junior engineer who has no memory of your project and knows only what is directly in front of it.

That one mental model changes everything about how you use these tools.

Why Context Is Everything

Every AI agent works within a context window — a fixed amount of information it can hold at once. Everything it knows about your project, your conversation, the files it has read — all of it lives in that window.

When the window fills up, older information gets pushed out. The agent forgets. It makes assumptions. It produces wrong output — not because it's broken, but because you didn't give it what it needed.

What this means practically:

  • The agent knows only what you show it
  • Ambiguity produces plausible but wrong guesses
  • Long sessions get progressively less accurate and more expensive
  • A well-structured context is your highest-leverage investment

The File Structure That Fixes 80% of Problems

The community standard is an AGENTS.md file for portable rules across all agents, with a CLAUDE.md that imports it for Claude-specific sessions. Here's a clean structure that works: marmelab

my-project/
│
├── AGENTS.md ← Shared rules (Codex, Copilot, all agents)
├── CLAUDE.md ← Claude-specific (imports AGENTS.md + extras)
│
├── .claude/
│ ├── rules/ ← Scoped rules per domain
│ └── commands/ ← Custom slash commands
│
├── docs/
│ ├── architecture.md ← How the system works
│ ├── database.md ← Schema and DB decisions
│ ├── authentication.md ← Auth flow
│ ├── api.md ← API conventions
│ ├── conventions.md ← Coding standards
│ ├── decisions.md ← Why things are the way they are
│ └── current-state.md ← Current progress snapshot
│
└── src/

Writing CLAUDE.md That Actually Works

Keep your root CLAUDE.md under 300 lines. Some high-performing teams keep theirs under 60. Every line costs context budget. Every unnecessary line dilutes what matters. buildcamp

What belongs in CLAUDE.md:

markdown

# Project: MyApp ## Stack - Backend: Node.js + Express + TypeScript - Database: PostgreSQL with Prisma ORM - Frontend: Next.js 14 (App Router) - Deploy: Vercel (frontend) + Render (backend) ## Key Commands - `npm run dev` — start local server - `npm test` — run test suite - `npm run lint` — lint check ## Architecture - Read `docs/architecture.md` before changing cross-cutting boundaries - API routes live in `src/routes/` - Follow patterns in existing route files ## Rules - Never modify the database schema without a migration file - All API responses follow the format in `docs/api.md` - Use existing error handling middleware — don't create new patterns

What does NOT belong:

  • Generic coding advice ("write clean code")
  • Long tutorials or changelogs
  • Task-specific notes that will be stale next week
  • Code style rules (use a linter for those)

CLAUDE.md instructions get followed about 70% of the time. Hooks enforce rules at 100%. For truly critical rules — never commit secrets, always run tests — use hooks, not instructions. datacamp

The Workflow That Actually Works

Step 1 — Explore first.
Before writing any code, ask the agent to read the relevant files and summarise its understanding. Fix misunderstandings here. It's cheap.

"Read src/routes/users.ts and src/middleware/auth.ts.
Summarise how authentication currently works."

Step 2 — Plan before implementing.
Ask for a step-by-step plan. Review it. Adjust it. Only then say go.

In Claude Code, /plan mode enforces this — the agent cannot write code until you approve the plan.

Step 3 — Implement in small steps.
One logical change at a time. Not "build the whole feature." One endpoint. One component. One migration.

Step 4 — Verify with real feedback loops.
Have it run your tests. Check the diff yourself. The agent cannot know if it succeeded without tests to run.

Step 5 — Commit after each working step.
Git is your undo button. Use it after every verified change so you can always roll back.

Writing Prompts That Get Good Results

The single biggest source of bad output is vague prompts.

Weak:

"Add authentication to my app."

Strong:

"Add JWT-based login to the Express API in src/routes/.
Follow the pattern in users.routes.ts.
Use the existing bcrypt dependency.
Add POST /login and middleware requireAuth.
Don't modify the database schema.
Done when npm test passes and I can log in with a seeded user."

The strong version has: goal, context, constraints, acceptance criteria, and scope. Every good task description needs all five.

Choosing the Right Model

Not every task needs the most powerful model. Using frontier models for simple edits wastes significant money.

Task Type

Model Tier

Why

Simple edits, extraction, classification

Small/fast (Haiku)

Low cost, sufficient capability

Daily development, feature work, refactoring

Mid-tier (Sonnet)

Strong coding at moderate cost

Architecture decisions, hard debugging, large migrations

Frontier (Opus+)

Deep reasoning for complex problems

Bulk mechanical tasks, subagents

Small/fast

High volume, cost-efficient

A common pattern: a stronger model plans, a mid-tier model implements, a small model handles bulk subtasks. This alone can cut costs by 60–70%.

Managing Context in Long Sessions

The highest-impact practice is clearing sessions at 60% capacity. When a session runs long, the agent starts working with degraded context. Starting fresh with a well-written prompt is almost always better than continuing a cluttered session. datacamp

Practical rules:

  • One task per session. Unrelated tasks in the same session pollute each other's context.
  • Point to specific files. "Read src/auth/jwt.ts" beats "search through the codebase."
  • Paste only the relevant error. The failing test and stack trace — not the entire log output.
  • Use subagents for heavy file reading so it doesn't consume your main context.

The Pitfalls That Cost People Hours

Hallucinated APIs. Agents invent functions that don't exist, reference outdated library versions, use packages that were deprecated. Always verify generated code against actual documentation.

Over-engineering. Agents add abstractions you didn't ask for. State "keep it minimal" explicitly when you need simple code.

Blind approval. Treat output as a pull request from a junior developer. Read every diff. You remain accountable for the code. eesel

Secrets in prompts. Never put API keys, passwords, or credentials in your prompts. Review generated code for auth flaws and injection vulnerabilities.

Arguing in circles. If two corrections fail, start a fresh session with a better prompt. Piling on context makes it worse, not better.

How Your Role Actually Changes

The agent types code. You do something harder.

You decide the architecture. You write the spec. You decompose the work. You review the output. You catch what the agent missed.

Your distributed systems knowledge, your system design instincts, your understanding of databases and networking — these are what let you evaluate whether the agent's output is actually correct. Without them, you're just approving code you don't understand.

The engineers who use AI agents best aren't the ones who prompt the most. They're the ones who think the clearest about what needs to be built before the agent touches anything.

A Starter CLAUDE.md You Can Use Today

markdown

# Project Context [Your project name and one-line description] ## Stack [Your actual stack] ## Key Commands [run, test, lint, build commands] ## Architecture - Read docs/architecture.md before cross-cutting changes - [Your key architectural boundaries] ## Rules - [Your 3-5 most important non-negotiables] - Run tests before marking anything done - Never modify [your critical files] without explicit instruction ## Docs - docs/api.md — API conventions - docs/decisions.md — Why things are built this way

Keep it short. Keep it factual. Update it every time you correct the agent on the same mistake twice — that's your signal that a rule belongs here.