AI Agent, Model & Token Usage Guidelines¶
- Default to medium reasoning and 200k context. Step up only for hard problems, then back down.
- Watch context size. Compact at around 60 to 70 percent.
- Keep the cache warm. Avoid long idle gaps mid-task, and compact before stepping away since the cache expires anyway.
- One task per session. Start fresh for unrelated work.
- Plan first. Agree on the approach and definition of done before any code is written.
- Tell the agent what you already know: files, root cause, test command. Discovery could be the expensive part.
- Keep a short instructions file per repo, such as
AGENTS.mdorCLAUDE.md, with architecture, commands, and conventions. - Use a cheaper model for boilerplate and mechanical work. Save the top model for design, hard bugs, and review.
- Avoid long autonomous runs unless the task is well-scoped and has automated checks the agent can run itself.
- Review agent PRs with the same rigor as human ones. Rework is the hidden cost.
- Track tokens per session to find waste, and cost per merged PR to judge value. If your coding environment does not show usage, find an extension or plugin that does.
Team setups¶
Collected via three questions: what's your setup, what do you use it for and when not, any learnings to share.
Summary¶
- Tools. Claude Code is the most common base, run via CLI, IDE extension, or T3 Code. Cursor is the IDE for most people, either for manual edits or running its own agent. VS Code for one. Codex is the usual second agent.
- Models. Opus 5.5 is the default for implementation. Fable 5.1 or a second provider such as ChatGPT or GPT 5.6 for planning, ambiguous problems, and cross-checking. A common pattern is plan with a strong model, implement with a cheaper one, then have the strong model review for bugs and edge cases.
- Effort. Most run high or xhigh reasoning by default, some run medium. Everyone drops effort for trivial tasks.
- Shared habits. Plan mode or a separate planning model for multi-file changes. Markdown files as the handoff between models and sessions, and as the upfront plan with phases, tasks, and checklists. When one agent is stuck, summarize the problem in a markdown file and hand it to another. Subagents for independent tasks in parallel. Manual review of the diff before pushing. Worktrees to avoid stash and pop.
- Frontend. Give the agent the full UX flow first, then have it build a Figma overlay on the site and iterate until nothing differs. A click-to-comment tool on the site lets the agent pick up visual feedback directly. ChatGPT, not Claude, for mockups, images, and visual design.
- Cost levers people found. Compact before heavy prompts. Skills for repeated expensive lookups such as database schema. Paste doc URLs instead of letting the model guess. Skip AI for constant changes and version bumps.
- Where people differ. Some build skills and memory files, others dropped them and let the model figure things out. Time spent thinking through what to build is now the main bottleneck, not execution.
- Common complaint. Claude replies are verbose regardless of instructions.
- Onboarding gap. Codex needs the Bedrock API key. Not obvious to new joiners.
Niklesh (25 Sep 2026)¶
Setup. T3 Code, a UI over the Claude Code, Codex, and Cursor CLIs. Easy context switching, cross-repo work, and worktrees. Cursor only for rare manual edits. Previously used the Claude and Codex extensions inside Cursor, but they were slow and would hang the laptop.
Models. Fable 5.1 was the go-to for planning, with implementation on Sonnet 5 or Grok 4.7. Recently moved planning and ambiguous tasks to Opus 5.5, which matches Fable for this. GPT models for cross-checking or documentation and wording.
Learnings. Plan first, write handoff docs between sessions, compact before context fills. Cross-check AI work at least at a high level.
Miquel (25 Sep 2026)¶
Setup. Cursor with Opus 4.6 on Bedrock, plus a personal ChatGPT Pro account.
Usage. All coding and review on Opus 4.6 in Cursor. For ambiguous work, narrows scope in ChatGPT Pro first since it tends to over-engineer, then iterates via a markdown file passed between ChatGPT Pro and Opus.
Learnings. Bounce big changes across two models. Ask Claude to read your recent PRs, learn from the review feedback, and store it in a markdown file.
Shubham (25 Sep 2026)¶
Setup. Claude Code only, via CLI in iTerm or the JetBrains extension. Opus 5.5 on xhigh for most coding, low or medium for trivial tasks. Several MCPs connected to guide exploration.
Usage. Almost everything including debugging. Plan mode for multi-file changes. IDE refactoring for constant changes and version bumps. Mistral and Gemini web for research and learning.
Learnings. Compact before a prompt that will spawn agents or heavy reasoning. Git worktrees to avoid stash and pop.
Nicola (26 Sep 2026)¶
Setup. Cursor Pro as IDE, with Claude Code and Codex on Bedrock as plugins. Considering VS Code since Cursor features are rarely used. Claude Code is the main assistant, Opus 5.5 and Fable 5.1 on high or extra high. Built skills for database schema and shot data lookups to cut repeated discovery cost.
Usage. Linear MCP to collect issues. For conceptual problems, brainstorms with Fable on extra high in a scratch folder with small web apps or Python notebooks to visualize the data. Hands integration to Opus with instructions not to commit, then reviews the diff manually. Codex with GPT 5.6 Sol on high as a second opinion. Pastes the automated PR review into Fable to triage which points are valid.
Learnings. - Multi-repo IDE workspace, such as studio-frontend plus studio-backend, so one session has cross-repo context and can commit to both. - Turn anything the model spends minutes figuring out into a skill if you will need it again. - Paste external doc URLs so the model works from current information. - Codex must be set up with the Bedrock API key. Not obvious at first.
Sravan (29 Sep 2026)¶
Setup. Cursor with Opus 5.5 and a 300K context window. Nothing done manually.
Usage. - Figma to code: gives the agent the complete UX flow first. The agent goes through Figma across all screen sizes and creates a task sheet for every screen and component. - Pixel-perfect work: the agent builds a Figma overlay on the site, compares the implementation against Figma, and iterates until there are no visual conflicts. - Feedback: a comment tool where Cmd + click anywhere on the site leaves a comment. The agent picks up the comments and fixes them. - The whole implementation is planned upfront in markdown files with phases, tasks, dependencies, and checklists. Subagents handle independent tasks in parallel.
Learnings. This has been the fastest workflow found so far.
Hemant (29 Sep 2026)¶
Setup. Cursor for faster output that is easier to understand in one read, plus the Claude Code extension on Cursor usage, usually Opus models.
Usage. Gets the plan sorted with a higher model, using skills for the heavy lifting. Writes code with a mid-tier model. Then asks a higher model to check the approach, edge cases, and bugs.
Gerard (29 Sep 2026)¶
Setup. VS Code with the Claude and Codex extensions. 95 percent of coding with Claude, lately Opus 5, before that 4.8. Fable rarely. ChatGPT 5.6 in the browser for planning, sketching ideas, and discussing frameworks or architecture, with projects for structure and shared memory.
Usage. Claude for coding, sometimes ChatGPT to review the implementation. When Claude is stuck, asks it to summarize the problem in a markdown file and passes it to Codex with GPT 5.6 Sol, which reliably solves it. ChatGPT for image generation such as mockups, UX, and creatives. Never Claude for design, images, or dashboards because of its style.
Learnings. - Execution is no longer the barrier, thinking is. Shape the idea clearly first so the model has fewer ways to build the wrong thing. - Stopped using skills and memory files. Models are good enough that these can limit them, and they often find better approaches on their own. - Claude is unnecessarily verbose and convoluted when answering questions, and drifts back to it regardless of rules.