You've read the Cursor guide, watched the Claude Code demo, did the math in your head, and decided: it's time. You send down orders to the team—try it next sprint. Two weeks later, you look at the numbers and see that lead time hasn't shortened, it's increased. Strange incidents start flying into the tracker. Two of your best developers walk around with faces that say 'I told you so.' At retro, you hear a measured 'we need more time to assess the impact.' In reality, that means 'get rid of this thing.'
A channel with guides and content on claude code, we post news (when they cut rate limits by 10x) and what tools we implement through claude for projects, channel: https://t.me/claudedevolper
Sound familiar? This is a typical picture of AI rollout in an engineering team through top-down administration. The problem isn't the tool, isn't the models, and isn't the skeptics. The problem is that a push model (forcing adoption) systematically doesn't work with senior developers—and the stronger your team, the worse it works.
This guide covers a pull model (drawing in, engaging). How to build things so that seniors choose to work with an agent of their own accord, and three months later become evangelists. This isn't about motivational speeches or bonuses for AI-generated code percentage. It's about an engineering solution: workflow, infrastructure, and deployment phases that pass the filter of an experienced developer.
Why Forced Adoption Fails
The temptation is understandable. The topic is hot, the board is asking questions, there are a couple of juniors in the team with stars in their eyes saying 'productivity tripled.' The logical step is to standardize. Deploy the tool to everyone, introduce a metric, report results at quarter-end.
Then comes what nobody expected.
Code review becomes a bottleneck. PR volume went up, reviewer load hit the ceiling. Generated code looks plausible, but it's 'heavier' and reads slower than handwritten: multi-layered constructs, convention mismatches, unexpected side effects in places you don't expect them. Seniors spend more time on review than juniors saved on writing. Net negative on the balance sheet.
Release cycles get longer. Entropy of approaches in the codebase increases—one style in one place, another style in another, a third somewhere else, all in the same repo. Architectural decisions are made without end-to-end visibility, because everyone has their own agent and their own context. After a couple of sprints, technical debt accumulates that nobody explicitly invested in—it just appeared. Meanwhile, incident count rises: some bugs slip through overloaded reviews.
Your best people quietly sabotage. Not openly—they just keep writing by hand, say in meetings 'I tried it, it didn't work for me,' react skeptically in chats to juniors' enthusiasm. This isn't stubbornness or Luddism. Under this are four concrete fears, each one a rational defense mechanism:
- Conflict with professional identity. A senior is someone whose value is built on the quality of their decisions. When you offer them 'describe the task—accept the generated code,' you're giving them the role of a technical scribe-reviewer, and giving the author role to the model. This isn't a role promotion, no matter how beautifully it's packaged in slides. It's a demotion.
- Fear of deskilling. Not imaginary, real. If you spend three years writing code mostly through generation, what happens to the skill of holding architecture in your head, feeling code smells, telling a right solution from one that just looks right?
- Loss of autonomy. The decision came from above, the tool was imposed, a metric for 'AI-generated code percentage' hangs over your head. A complete inversion of the mode strong engineers have historically worked in—and the very mode from which their expertise grew.
- Threat to job security. Not 'the model will replace me'—that's a junior's fear. The senior version is more subtle: 'if my value boils down to reviewing generations, I'm interchangeable.' The more solid a senior's standing, the more acutely they feel it.
The push model tries to ignore all four points and force through with KPIs. The pull model leans on them as design requirements. That's all the difference.
What a Senior Really Wants from AI
If you flip the four fears from the previous section 180 degrees, you get a requirements list. Not for the tool—for the work mode. These aren't 'features' you can buy with a subscription. These are workflow properties that either exist or don't.
A senior wants to remain the author of decisions. Not a reviewer, not an operator managing the model, not a recipient of generated code. The author. This means the agent is engaged before an architectural decision is made and helps make it, rather than implementing one already made. Paradoxically, it's exactly the early entry of the agent—at the exploration and brainstorming phase—that removes the fear of 'I'm reduced to reviews,' because the author role remains with the human in this setup. During exploration, the agent is a sounding board. During implementation, the agent is hands. Not the other way around.
A senior wants their expertise to strengthen, not atrophy. This is a subtle point. Any strong engineer knows: skills grow where you have to think, and atrophy where you stop thinking. If the agent takes on 'thinking,' the skill goes away. If the agent takes on 'searching, verifying, synthesizing routine,' while thinking stays with the human—skills grow faster than without the agent. The workflow must explicitly separate these two modes. Not 'the agent did everything, I just clicked OK,' but 'I guided the agent through my thinking, and the output became the code.'
A senior wants personal control over the process. Not 'try the new tool by Monday,' but 'here's a capability, explore it at your pace.' Optionality here isn't politeness—it's a design requirement. If the agent is forced on you, any of its failures go into the 'I told you so' pile. If you chose it, the same failures are seen as the cost of your own decision, not a system failure.
A senior wants their role to become less interchangeable, not more. This is the most counterintuitive part. The push model implicitly says: 'you're all AI operators now, the differences between you have blurred.' The pull model should say the opposite: 'after adoption, you're even more valuable, because working effectively with an agent is a separate skill, and you have it while juniors don't.' Paradoxically, that's exactly how reality works: an effective engineer's ROI from working with an agent is an order of magnitude higher than a weaker one's. The strong engineer asks the right questions, spots flaws in the output, maintains scope boundaries, keeps the agent from going down rabbit holes. The weak one just generates garbage faster than before.
These four points aren't marketing. This is the spec for the rest of this article. Next, I break down the principle, workflow, infrastructure, and deployment phases—and in each section explicitly show which of these four points each mechanism addresses. If your future rollout leaves one of the four uncovered, it won't pass a senior's filter, no matter how beautifully the rest is packaged.
The Core Principle: Agent in a Subordinate Proactive Position
I repeat this formula constantly, and every time I see the question in people's eyes: 'wait, how? proactive means doing it on its own, right?' No. This is the key to everything.
Most demos and blog posts about agentic development are built on the opposite model: you give the agent a task, it decomposes it on its own, implements on its own, opens the PR on its own. The human formulates at the start, accepts at the end. Let's call this the autonomous position. It looks good on Twitter, works poorly on a real codebase, and categorically doesn't pass a senior's filter. Because it requires trust that a senior doesn't have—and has nowhere to get it from.
Subordinate proactive is a different mode. Let me break it down on two axes separately.
Subordinate means the agent makes no decisions. None. It doesn't choose architecture, doesn't approve scope, doesn't decide if the neighboring module needs refactoring, doesn't open a PR without explicit command. All decisions are on the human. This gives the senior that personal control that doesn't exist in autonomous mode. And, more importantly, preserves their authorship—every decision in the code is actually theirs, not the model's.
Proactive means the agent doesn't wait to be asked. It runs through the codebase on its own, checks hypotheses on its own, raises contradictions, suggests alternatives, and reminds you what you missed. Not 'waiting for a prompt'—but 'I see a risk emerging here, what do we do?'
In practice from real life—this is how a senior developer works with a strong junior intern who has deep knowledge, instant speed, and zero responsibility. The intern is proactive—asks questions, suggests options, digs deeper than asked. But decisions are made by the lead. The intern doesn't push to main without approval and review, doesn't rewrite neighboring modules 'while at it,' doesn't think they know better. They amplify the lead's thinking process, don't replace it.
This analogy works not just for understanding. It works as the foundation for designing the work process without inventing approaches from some other universe.
Here, the main deskilling fear is addressed. When the agent is in a subordinate position, thinking continues to be done by the human. The agent takes on searching the codebase, verifying hypotheses, synthesizing alternatives—the routine that doesn't build skills. But thinking—where to draw the scope boundary, what trade-off is acceptable, what abstraction is right—stays with the engineer. After half a year in this mode, skills don't atrophy, they grow faster, because the engineer processes more tasks in the same time, and each task requires exactly thinking, not search busywork.
A second property follows from subordinate proactivity, one often overlooked: the agent must be engaged early in the process. Before the first line of code is written. Before you yourself know the exact answer. At the stage of 'I understand in general what needs to be done, but I'm not sure about the details.' This is exactly where the agent delivers maximum value—because it runs through the codebase faster than you, holds more context at once, and doesn't get tired from the twentieth clarifying question.
If you bring in the agent late—at the stage of 'write me a function that does X'—you're using it as a code generator. That's exactly the mode that causes senior rejection. Because in it, the agent truly replaces the author instead of amplifying them.
The subordinate proactive position isn't a setting. It's a construct made of three things: workflow, instructions (a context layer), and developer habits. Workflow sets the rhythm—where the agent enters and where it stops. Instructions set boundaries—what it can and can't do. Habits set quality—how well you can talk to it. Next, we review all three in order.
Workflow
By itself, the subordinate proactive position is just a concept. Without a process, it doesn't materialize. Workflow is where the concept turns into daily practice, and it's also the main artifact you as a tech lead need to either create or take ready-made and adapt.
The process is simple: instead of declarative commands ('write a function,' 'find the bug,' 'generate tests'), the agent and engineer move in an interactive cycle—explore the task, check hypotheses, discuss options, lock down spec and acceptance criteria, decompose, implement in phases with the ability to intervene without breaking the process, run checkpoints with intermediate reviews, cover with tests, and finally review before PR.
This is not just a nice sequence. These are three phases with different roles for the agent in each.
Brainstorming. The most important phase. Here the agent is a discussion partner, not an executor. It receives intent as "we're working on feature X, here are user stories, here are acceptance criteria." Next comes Q&A mode. The agent researches the codebase, highlights contradictions, asks clarifying questions, and proposes alternatives. The output is a design document with a fixed structure: goal, scope, non-goals, architectural principles, contracts, risks, and decision log.
Here we address the fear of "being reduced to code review." The author of all decisions in the design document is the human. The agent is a tool for structuring their thinking, not a replacement.
Planning. Once the design is fixed, we switch to "how." The goal is to break the task into sprints, each with its own scope, acceptance criteria, and blast radius assessment. Sprint = an atomic unit of work that doesn't leave the code in a broken state. At this phase, the agent goes through the code again, but the focus shifts from "what to build" to "where to put it and what it will affect." Often at this stage, nuances missed during brainstorming surface.
Here the senior dev remains the architect of the implementation. They don't write the plan independently—but they approve each turn. This is precisely that autonomy, preservation of self-worth through expertise and decision-making authority.
Work cycle. The cycle plan → implement → review → fix → review → commit, repeated per sprint. Each sprint gets its own session. Between sprints—commit is mandatory. Below is the diagram.

Separately on review. Review strictly in a separate session. Not in the same one where implementation happened. In one session, the model is biased toward its own results—just like us, to be honest. In a separate one—it catches what the biased one missed. Personally, besides visual review, I run two sessions with different models in parallel: Claude and Codex, for example. They find different things, and it objectively improves quality.
Here, by the way, the main new skill for engineers in the era of agentic development crystallizes—precision, clarity of thought, system thinking, and patience. The cycle "human reviews → agent fixes" works exactly as well as you can explain what's wrong. If you ramble—the agent will go off the rails. If you answer crudely and off-topic—you'll get a workaround. It's just like working with a real person, except you can't fire them.
Skills—the thing without which nothing works
The described workflow is impossible on a bare model. Each phase requires specific behavior from the agent that no frontier model demonstrates on its own. Under normal instructions, it will write code—but we need it to first ask questions, then explore the codebase, then fix decisions in a file, and only then write code.
This is solved with skills—separate files with contextual instructions that load when the agent needs to enter a specific mode. Technically—modified system prompts, but in effect—like changing roles. With skills, models perform miracles that they usually don't pull off with bare instructions.
Minimal working set—three skills:
- brainstorm — switches the agent into discussion partner mode. Takes intent, researches, challenges, locks in decisions, outputs a design document.
- planner — switches into implementation architect mode. Takes a design document, validates it against real code, breaks it into sprints, outputs a development plan.
- code-review — switches into reviewer mode. Takes the sprint implementation result in a fresh session, outputs a report.
I've published the skills in open access: github.com/pridees/skillforce. They're tuned to work together, tested on Claude Code, Codex, Gemini, Open Code, and several other harnesses, on models from Anthropic, OpenAI, Kimi, and GLM. No guarantees it'll work identically for you—models differ, harnesses differ, codebases differ—but as a starting point they work fine. Then you adapt them to your specifics.
Installation (may require Node.js installed):
npx skills add pridees/skillforceОбъяснить с
As a team lead, you shouldn't force each senior dev to assemble this from scratch. One of your key contributions—prepare skills and the contextual layer centrally, in a format convenient for attaching to the repository. This eliminates a huge chunk of the "onboarding tax" for senior devs who otherwise would never get to configuration independently. And at the same time—it removes the diversity of approaches in the team that we saw in the push-scenario.
What should appear in the repository
A team lead doesn't need to know every prompt—but must know what should appear at the output of each phase. This is your main tool for checking that the workflow works, not simulates.
Design document (after brainstorm). Structure: Understanding Summary, Non-Goals, Assumptions, Design Principles, Data Model / Contract, Runtime Behavior, Testing Strategy, Decision Log, Acceptance Criteria.
Development plan (after planning). Structure: Overview, Prerequisites, Sprint 1...N (each with its own goal, tasks, acceptance criteria, validation), Testing Strategy, Risks & Rollback.
You can keep them in the repo; I usually delete them a couple weeks after release. If a bug is found in the code later, the commit hash, spec, and plan help restore context for debugging faster.
Infrastructure setup
Skills are half the job. The other half—what sits in the repository and loads when the agent works. It's the harness that everything spins on, and the contextual layer that determines how the agent understands your project.
Harness
Harness is a layer between "intelligence" (the model) and "execution" (your code and tests). It's essentially "the agent" in the colloquial sense—Claude Code, Codex CLI, Cursor agent mode, Open Code. Their job isn't to "write code instead of the developer," but to close a managed SWE cycle: task intake, codebase exploration, planning, making changes, instruction sync, tool and MCP calls.
The key requirement—ability to check itself. Inside the agent cycle there must be a stage of self-criticism and self-checking: running compiler, linter, tests. Or the ability to configure this manually via hooks. Without this, harness is not an SWE tool, just a pretty chat with autocomplete.
There isn't much choice. Mainstream options—Claude Code and Codex CLI, both can switch to third-party provider models (if that's politically acceptable in your company). If not—my personal recommendation is Pi, an extremely extensible thing that lets you add everything you need. For models—take the best available; if frontier is unavailable, in open weights right now models at the level of GLM/Kimi/MiniMax pull well.
As a team lead, you pick a harness once for the whole team and lock it in as standard. Heterogeneity here is evil: different harnesses read different instruction formats, and a contextual layer tuned for one doesn't work with another.
Contextual layer
This is the heart of your infrastructure. Here live all the agreements, conventions, guidelines, skills, and patterns that distinguish your codebase from any other. What you transmit to a new developer over the first two months, you need to transmit to the agent through files.
I build this layer like this:
/ ├── .agents/ │ ├── rules/ │ │ ├── coding-guidelines.md │ │ └── conventions.md │ ├── skills/ │ │ ├── brainstorm/SKILL.md │ │ └── code-review/SKILL.md │ └── AGENTS.md Объяснить с
Next—symlinks for the specific harness: .agents/AGENTS.md → .claude/CLAUDE.md or .agents/AGENTS.md → .codex/AGENTS.md. OSS agents understand .agents out of the box or allow you to specify the name in config.
AGENTS.md—the main agent instruction. Basic structure: project definition, development rules (with links to rules/), repository structure, spec guideline, commit conventions, red flags (what we never do), build and test commands, documentation links. A minimal universal template to start from, I posted separately as a gist—gist.github.com/pridees/82eef0e1710196188492695baef20ee6. Take it, adapt it to your stack.
Substantively, AGENTS.md isn't filled manually from scratch. There's an easier way: ask the agent to introspect the repository and its rules and update @.agents/AGENTS.md, keeping the structure. Next—review and edit. Once, for the whole team.
Main principle: "the model will figure it out"
This is a counterintuitive thing, usually understood only after a few iterations.
The temptation when setting up—describe every case for every eventuality. All rules, all conventions, all exceptions. Don't do that. Here's why:
- Bloated AGENTS.md takes up space in the context window. Every instruction token is a token stolen from the task. The more rules, the worse the agent solves.
- Instructions conflict with real code. Reality is always richer than written description. The agent sees code—and predictably writes as it sees. And it irritates you that it "doesn't follow the rules."
- Imperative rules are fragile. "Write controllers like this"—but what if the controller is different? "Don't use X"—but what if one place needs it? Every exception requires instruction updates.
A different approach works much better: clarify details as you work with the agent, and make your code the source of examples. If you have a well-designed module or a successful feature implementation—just describe the case and point to these files. "When working with the domain layer, look at src/domain/order as an example." Instead of a page of instructions—one line plus a real example from your own codebase.
Same with rules: don't write You MUST read for every file in rules/. Instructions about REST controllers aren't needed for domain layer work, and vice versa. Just describe the condition when the instruction is relevant—the model will figure it out.
Repeating patterns and prompts that keep cropping up for you or developers—pull them into separate skills. The contextual layer is alive; it grows with your team.
As a team lead, your task is not to write a perfect AGENTS.md once, but to create a process for its evolution. One person (you or a designated maintainer) keeps the contextual layer current, accepts PRs on rules and skills from the team, tracks what's started to repeat. This isn't a one-time activity; it's a new ritual in the repository—like CODEOWNERS, but for context.
Rolling out to the team
Harness chosen, contextual layer assembled, skills in the repo. Next comes the hard part. Turn infrastructure into practice without sliding back into the push-model we just deconstructed.
I break the rollout into three phases. Each has its own logic, metrics, and ways to mess it up.
Phase 1. Early adopters
Find 1–2 people in the team who actually want to work with this. Don't persuade, don't "give it a chance." If they want it—they've already tried it on the side, grown disillusioned with naive approaches, and are ready to put effort into a proper workflow. The ideal candidate is a senior dev with mature skepticism, not a junior with burning eyes.
Their task at this phase is not to "migrate the team to AI". Their task is to live through the workflow on real tasks and produce visible artifacts: design documents, development plans, clean PRs that are reviewed quickly and don't require rewriting. Several features brought to production. That will be your proof.
Don't set a deadline. Don't introduce metrics. Don't report to the board on this phase. Any publicity breaks it — volunteers start working for the report, not for quality.
Phase 2. Expansion
When the guys get credible results — you open an invitation. Not "now everyone will try it", but "here's what we tested — now it's available to anyone who wants it". Skills and context layer in the repository, documentation, maybe one internal meetup where colleagues show from experience how they work. Without pressure.
Next, you observe. Someone will join right away, someone in a month, someone in three. That's normal. Pull-model isn't about everyone joining on the same day. Pull is about them coming on their own.
Don't touch the skeptics at all. Some of them will join after seeing their colleagues' results. Some — after they hit the wall that the workflow solves. Some — never. You need to be able to live with the last scenario too: not every developer should work with agents, and trying to force the workflow on everyone is that same push-model in a new wrapper.
Phase 3. Standard
When 70–80% of the team consistently uses the workflow — you formalize it. Not before. AGENTS.md in the repo is mandatory, harness is standardized, skills are centralized and maintained. At this phase — yes, you can require it. But the requirement doesn't sound like "use AI", it sounds like "follow the repository conventions". If someone codes by hand — fine, the main thing is the conventions.
Never introduce a KPI "percentage of code from AI". This is the most popular and most destructive metric in all push implementations I've seen. It optimizes exactly what shouldn't be optimized and encourages exactly the behavior the workflow is meant to protect against.
What about metrics:
- Lead time from picking up a task to merging a PR. Should not grow. If you see a speed boost right off the bat — the team relied too much on AI, which is also not good. Improvement in this metric should be considered together with the metrics mentioned below.
- Review Throughput per reviewer — how many PRs per week per reviewer and how much time is spent. Your canary against a hidden return to push-model. If someone on the team starts generating in "made it — sent it" mode (the very behavior we're moving away from), reviewer load will spike before other metrics catch it.
- PR Rework Rate — the share of PRs that require significant changes after the first review (not cosmetics, not comments, but logic rewriting). Should not grow. If it grows — it means brainstorm/planning are being skipped, and the workflow degrades into "generation on demand".
- Mean Time to Recovery — time from incident to recovery. Indirectly shows how well the team understands the code it produces. Here's the most insidious risk of AI adoption: code is written faster, but if the author doesn't fully understand it (because they delegated more than they thought) — it will surface in an incident. MTTR grows = deskilling happened, need to review the workflow toward greater human involvement in the brainstorm/planning phase.
- A simple poll within the team every sprint or two. One question: "how much does the workflow help or hinder your work?". NPS format. The number doesn't matter — the trend does.
What NOT to measure:
- Percentage of code written by AI. Meaningless metric, trivially optimized, encourages harm.
- Generation speed. Measures the tool, not productivity.
- Volume of lines of code in a PR. The agent tends to bloat solutions, this metric encourages exactly that.
Harmful antipatterns
A short checklist of what breaks pull-model in one step:
- Deadline for implementation. "By the end of the quarter everyone uses it" = push.
- KPI "percentage of code from AI". Already explained.
- Public comparisons. "Look, Vasya did it 2x faster because he uses an agent" — guarantees Vasya becomes an outcast and the rest start deliberately sabotaging.
- Banning manual writing. If the task is small or obvious — workflow is overkill. Ban = pointless frustration.
- Forcing early adopters to teach others. Their job is to produce results, not be evangelists. Evangelism should be organic, otherwise it repels.
- Accepting metrics from outside. If implementation standards come from above your team — you don't have a team, you have a shop. Protect autonomy.
Your contribution as a team lead here is not in the speed of implementation. Within 3–6 months, a mature engineering practice should grow in your team, not a burnt-out collective in a state of mutiny.
Finale
A few months after implementation, you look back and realize something strange. The most valuable result wasn't an increased number of closed tasks, not reduced lead time, and not even satisfied seniors. The most valuable result was how thinking changed in the team.
For an agent in a subordinate proactive position to deliver what's needed — we had to learn to articulate intent more precisely than before. Dig deeper into the question of "what" we're building and not take our eyes off "how". Decompose the task before the first line is written. See the scope and fix it in words. State arguments in review so as not to get a workaround on the next iteration. Patiently go through the cycle instead of "I'll do it faster myself right now".
These skills — precision, clarity of thought, systematicity, and patience — are the main gain. They weren't given as a gift by the tool. They grew because the workflow demanded them every day. Artificial intelligence is already good enough to solve our everyday tasks, but it brightly highlights what we often swept under the rug — how we convey our thoughts and intentions to each other.
And that's exactly why push-model will never give this result, no matter how many KPIs you introduce. Coercion strips the engineer of authorship and replaces thinking with generation. Engagement — makes thinking work harder than it did without an agent. The difference isn't in speed. The difference is what happens to the team in a year.
Your role as a team lead is to create a space where this is possible. Skills, context layer, deployment phases, protection of team autonomy from external metrics. The seniors will handle the rest themselves. That's what they're there for.
Channel with guides and content on claude code, we post news (when they cut limits 10x) and what tools we implement through claude for projects, channel: https://t.me/claudedevolper
