AI-assisted software development did not replace our engineering team. It changed what being an engineer means on our team.
At the start of 2025, we were a small, distributed startup team: around seven people across engineering, product, and design functions, spread across the United States, England, Brazil, Spain, and Ukraine. We had no customers yet, no spare capacity, and a deadline that always seemed one week too close.Like most early-stage teams, we wore too many hats. People worked beyond their original specialisms because the product needed it. When someone was sick, on holiday, or simply working in a different time zone, progress could slow down quickly. We had code completion, bug-fix suggestions, and the occasional AI-assisted experiment. But we still thought of AI as a clever autocomplete tool, not a new way to build.
Then one of our team leads made a call that split the room: let AI help build the application, and let us guide it.
We were sceptical. Engineers tend to be. Six months later, we had an MVP in production across two regions, three web applications, one mobile application, multiple integrations, and supporting microservices—with the same team size. We now spend far less time writing first-draft code and far more time specifying, planning, reviewing, testing, and improving it.
This is not a story about handing software development to a model and walking away. It is a story about building a better operating system for engineering: one where documentation is the source of truth, small tasks protect quality, automation speeds up routine work, and people remain accountable for every decision that matters.
AI is only as useful as the context you give it, and only as safe as the humans reviewing it.
Before we changed tools, our main constraint was not a lack of ideas. It was the cost of coordination.
In a small distributed team, a feature can require product context, design decisions, backend knowledge, frontend execution, data changes, integrations, testing, and deployment. If this knowledge lives partly in people’s heads and partly across chat threads, a team loses time every time it needs to reconstruct the context.That cost compounds when people are working across time zones. A question asked at the end of one person’s day may not be answered until the next. A specialist may be unavailable. A feature can stall because the person who understands one small corner of the system is away.
We did not want AI to replace expertise. We wanted it to help make expertise more available, more structured, and easier to apply across the product.That distinction is central to AI-assisted software development. The goal is not “generate more code.” The goal is to reduce the friction between a well-understood problem and a well-tested solution.
| Before | What created friction | What we needed instead |
| Knowledge was distributed across people, repositories, and conversations | Context had to be rebuilt repeatedly | A reliable, living source of truth |
| Work arrived as broad feature requests | Scope expanded and assumptions multiplied | Small, testable slices of work |
| Specialists became bottlenecks | Progress depended on individual availability | Better shared understanding and review |
| Reviews focused on what had already been written | Errors appeared late | Planning, verification, and feedback earlier in the cycle |
| AI was treated as a code-completion tool | Its value stayed limited to isolated tasks | An engineering workflow designed around context and oversight |
Our first serious experiments came from a practical need. We did not have a designer at one point, but we had design demands. A team lead tested AI tools to explore ideas faster than she could have alone, especially outside her formal domain. The output was not perfect. But it was fast enough to be useful, structured enough to critique, and flexible enough to iterate.
That changed the conversation.The important breakthrough was not simply that models could produce more code. It was that they could work with broader project context. With the right instructions, reference materials, and tools, an agent could read patterns in a repository, follow team conventions, and help move a task through several steps instead of producing one isolated snippet.
Anthropic describes Claude Code as an agentic environment that can read files, run commands, make changes, and work through problems; its documentation also emphasises that effective work depends on managing context deliberately . GitHub’s guidance makes the same practical point from another toolset: break complex tasks down, be specific about requirements, provide relevant context, and validate the resulting code [2].
Those principles matched what we were learning in real time.We did not need AI to know everything. We needed it to know enough about this product, this stack, this set of conventions, and this particular task to make a useful contribution.
The most important lesson from our shift to AI-assisted software development is almost disappointingly unglamorous: documentation is the base of everything.
Before AI, weak documentation was inconvenient. A strong engineer could often compensate by asking the right colleague, reading enough code, or experimenting until the system made sense. With AI, weak documentation becomes a multiplier of risk. If the agent has incomplete or outdated context, it may fill the gaps with plausible assumptions. That is where confident but incorrect output begins.So we started treating documentation as a product artifact, not a side task.
We documented our stack, patterns, architecture, workflows, conventions, and recurring decisions. We divided the application into manageable pieces. We made it clear what “done” meant before asking a tool to build anything.
Claude Code’s documentation describes CLAUDE.md files as persistent project instructions for conventions, architecture, build commands, and common workflows. It recommends keeping those instructions concise, specific, and shared through version control so a team can refine them over time [3]. Whether a team uses Claude Code, another agentic tool, or a different internal system, the operating principle is portable: write down what you would otherwise have to explain repeatedly.
| Weak context | Strong context |
| “Build a new onboarding flow.” | “Build step two of the onboarding flow using the existing form pattern, validation standard, API contract, and acceptance tests.” |
| Generic coding conventions | Repository-level conventions, architecture decisions, test commands, and known constraints |
| Long, unstructured instructions | Concise documentation with clear sections and scoped rules |
| Knowledge held by a few specialists | Shared, maintained project context available to the whole team |
| AI guesses when information is missing | AI can retrieve the relevant rules, files, and examples before acting |
We also learned that documentation must be revisited continuously. It is not a “write once” asset. Every change in the product, architecture, or workflow is a chance to strengthen—or weaken—the context that future work depends on.
We did not eliminate our software-development lifecycle. We made it more explicit.Our practical loop now follows six stages:
| Stage | What happens | What AI can help with | What the team owns |
| Docs | Define the problem, constraints, architecture, and acceptance criteria | Retrieve relevant context and identify missing details | Confirm the source of truth is accurate |
| Plan | Break work into small, reviewable steps | Draft implementation plans and identify dependencies | Choose scope, sequencing, and trade-offs |
| Code | Implement one manageable slice at a time | Generate initial code, tests, migrations, or refactors | Review logic, maintainability, and alignment with specs |
| Test | Run automated checks and reproduce edge cases | Suggest test cases and help interpret failures | Set verification criteria and judge coverage |
| QA | Validate behaviour, design, and user experience | Surface inconsistencies or support exploratory checks | Confirm the product behaves as intended |
| Deploy | Ship through controlled release processes | Assist with documentation and deployment preparation | Approve, monitor, and own production outcomes |
The most important rule is what happens between stages: we go back to the documentation.If a plan exposes a missing decision, we update the docs. If code reveals a constraint, we update the docs. If QA uncovers a gap between the specification and the real experience, we update the docs. The workflow is not linear because software is not linear. But documentation gives the team a stable point to return to.
This resembles the workflow Anthropic recommends for larger coding tasks: explore first, create a plan, implement, then verify the result against explicit checks [4]. It also mirrors GitHub’s advice to provide specific requirements, break down complex tasks, and use tests, linting, and security tooling to validate AI-generated output [2].
The tools are evolving quickly. The discipline is more durable.
There is a temptation to hand an AI agent a large feature and ask it to “build the whole thing.” We have learned—sometimes the hard way—that this is where complexity becomes dangerous.
A multi-step wizard, for example, may appear to be one feature. In practice, it contains structure, state management, validation, error handling, user feedback, API interactions, accessibility requirements, analytics, tests, and deployment considerations. Asking for all of it at once can create output that looks coherent but contains hidden assumptions.We now prefer a more deliberate approach:
1. Establish the structure.
2. Define one step.
3. Implement and test that step.
4. Review it against the specification.
5. Update context if necessary.
6. Move to the next step.
This is not slower. It is usually faster because it reduces rework.
Claude Code’s best-practice documentation explains why: context is finite, and performance can degrade as a session becomes overloaded. It recommends separating exploration and planning from implementation so an agent does not solve the wrong problem [4]. In other words, a smaller task is not merely easier for a model. It is easier for a human team to review, test, understand, and reverse if needed.
Small pieces are not a compromise. They are a quality-control mechanism.
The role shift has been real.We are not saying engineers no longer write code. We do. But the centre of gravity has changed. More of our value now comes from clarifying the problem, defining constraints, setting acceptance criteria, reviewing generated work, testing behaviour, and making decisions that require product and systems judgment.
That has broadened our team in useful ways. Backend engineers can prototype user interfaces. Frontend engineers can understand and apply data migrations. Product people can participate more directly in technical exploration. The old boundaries have not disappeared, but they have become more permeable.The skills that matter most are changing:
| Less valuable as a standalone skill | More valuable in AI-assisted software development |
| Producing boilerplate quickly | Turning ambiguity into an explicit specification |
| Knowing one layer of the stack in isolation | Understanding how decisions affect the system end to end |
| Handing off work with minimal context | Creating context another person—or an agent—can use reliably |
| Treating testing as a final checkpoint | Designing verification into the task from the beginning |
| Accepting generated output because it compiles | Reviewing for correctness, security, maintainability, and product fit |
This is not an argument against deep specialists. In fact, a team still needs people who can go deep where complexity demands it. But it is an argument for a more generalist operating model: one where people can contribute across boundaries while maintaining rigorous standards.
Model Context Protocol (MCP) became another meaningful step for us.MCP is an open-source standard that lets AI applications connect to external systems such as data sources, tools, and workflows [5]. In practice, that means an engineering agent can access the right kind of context or capability at the right time instead of relying only on what is inside a single chat or codebase.
For our team, MCP made it possible to build a more useful toolbox around the agent. One example is Context7, which we use to bring current technical documentation into the workflow without forcing the agent to rely on stale or generic assumptions. We also drew value from Superpowers, an open-source toolkit that helps structure planning and specification work for more complex solutions.
The benefit was not “more tools.” The benefit was better-grounded work.
But MCP changes the security equation too. The official MCP guidance recommends connecting only to trusted remote servers, reviewing permissions requested during authentication, and using tool-level permissions to control what an agent can do [3]. Its security guidance details the risks that can arise when tools and external systems are connected without sufficient authorisation and safeguards [7].
That means a mature AI-assisted engineering workflow needs two habits at once:
| Capability | Responsible practice |
| Connect tools and data sources | Use only trusted sources and make ownership clear |
| Give agents access to documentation | Limit access to what is needed for the task |
| Allow tools to take actions | Review permissions and separate read from write capabilities |
| Share repository context | Protect secrets, credentials, customer data, and sensitive configuration |
| Add new MCP servers | Assess trust, scope, maintenance, and security before connection |
The ability to connect an agent to more systems is powerful. The responsibility to understand those connections is non-negotiable.
We did not get everything right on the first try.
AI accelerated the creation of work. That meant alignment had to happen faster too.
In a team spread across five countries, communication was already difficult. With AI speeding up prototyping and implementation, the cost of misalignment increased. A vague instruction could now produce more output more quickly. That is not progress if the output moves in the wrong direction.
We had to create stronger rhythms for documenting decisions, reviewing plans, and clarifying ownership. The more autonomy an AI tool has, the more intentional a human team must be about direction.
AI capacity is not unlimited. Usage quotas, context windows, model selection, and cost are part of the operating reality.
We learned to use stronger models where the work demanded deeper reasoning or foundational documentation, and to avoid spending expensive capacity on tasks that could be handled with a smaller model or a simpler workflow. We also learned to keep prompts and tool outputs concise where possible. Tools such as Caveman helped us compress context in some scenarios.
The goal was never to optimise token use at the expense of quality. It was to treat AI capacity as we would any other constrained engineering resource: with intention.
At the beginning, we sometimes treated the system like a factory: give it all the documentation, ask for the entire feature, then expect a finished result. That approach failed.
Language models can produce output that is plausible but incorrect, incomplete, insecure, or misaligned with the actual system. GitHub explicitly advises developers to understand, review, and test AI-generated code; its responsible-use guidance notes that agentic outputs may contain inaccuracies, security issues, or hallucinated findings and must be validated by people before merging [8].
That is why every generated change still goes through review, tests, and QA. We do not outsource accountability. We use AI to increase the speed at which accountable people can work.
Our team size did not increase. But our ability to move did.Today, we have an MVP in production across the EU and US, with three web applications, one mobile application, multiple integrations, and microservices supporting the product. Those outcomes reflect many factors: strong people, changing priorities, hard work, and a willingness to learn. AI did not create the product alone.
What AI changed was our leverage.It helped us move from a team constrained by specialist availability and manual implementation capacity toward a team that can specify, explore, and iterate more broadly. We can test ideas earlier. We can expose ambiguity faster. We can share knowledge more deliberately. And we can spend more energy on the decisions that make the product useful.
| What changed | The practical result |
| Documentation became a shared engineering asset | Faster onboarding, fewer repeated explanations, and more consistent implementation |
| Plans became explicit before implementation | Smaller scope, clearer reviews, and less rework |
| AI became part of the delivery workflow | Faster initial drafts, tests, exploration, and iteration |
| Human review became more structured | Better alignment with product requirements, architecture, and quality standards |
| Team members worked across traditional boundaries | More resilience when capacity or availability changed |
| Deployment and checks became more automated | More repeatable releases and clearer evidence before shipping |
The reward is not just speed. It is resilience.
For teams considering a similar shift, start with this checklist.
| Question | Ready? |
| Do we have a maintained, readable source of truth for architecture, conventions, and test commands? | ☐ |
| Can we describe the user problem and acceptance criteria before asking an agent to implement it? | ☐ |
| Are tasks small enough for a human reviewer to understand and validate? | ☐ |
| Does every workflow include tests, quality checks, or observable verification criteria? | ☐ |
| Do we know which decisions must retain human judgment and accountability? | ☐ |
| Are secrets, customer data, and permissions protected when agents access tools or repositories? | ☐ |
| Do we trust each MCP server and understand its access scope? | ☐ |
| Do we measure quality, rework, and employee experience—not only output speed? | ☐ |
| Do we update documentation after meaningful product or technical changes? | ☐ |
| Do engineers have time to learn, review, and improve the workflow itself? | ☐ |
If several answers are “not yet,” that is not a reason to avoid AI. It is a signal to strengthen the base before increasing autonomy.
We build products intended to make work more human: helping employers understand what their people value, make benefits easier to access, and use workforce insight to improve retention.
That mission shapes how we think about engineering too. We do not believe technology should make people invisible. We believe it should remove the friction that prevents people from doing their best work.
For our engineering team, AI gives us more capacity to build, test, and improve. For the employers we work with, better technology should create more capacity to listen, support employees, and act on what matters. In both cases, the principle is the same: automation should increase human value, not reduce it.
Six months ago, we were a small team wearing too many hats and feeling the strain when a deadline tightened or someone was unavailable. Today, we are shipping a product across two regions with the same headcount and a very different way of working.
AI did not make us faster at typing code. It changed how we approach engineering.The teams that get the most from AI-assisted software development will not be the ones that hand work to a model and walk away. They will be the ones that invest in context, keep tasks manageable, verify relentlessly, connect tools responsibly, and continue to put people in charge of the decisions that matter.
Build the base. Guard the base. Then let AI multiply your discipline as much as your speed.If this sounds like the kind of engineering culture you want to help build, explore careers at SideUp.
AI-assisted software development uses AI tools to support work such as planning, code generation, testing, documentation, debugging, code review, and deployment preparation. Engineers remain responsible for requirements, technical decisions, review, quality, security, and production outcomes.
AI can accelerate implementation, but production software still needs human ownership. Engineers must define requirements, provide accurate context, review generated changes, run tests, validate security, and make accountable decisions about deployment.
Documentation gives AI tools the context they need to follow a project’s architecture, conventions, workflows, and constraints. Without reliable context, an AI tool is more likely to make plausible but incorrect assumptions.
Model Context Protocol (MCP) is an open-source standard that enables AI applications to connect to external data sources, tools, and workflows. It can make AI agents more capable, but teams should use trusted servers, review permissions, and protect sensitive systems and data.
Teams should apply the same—or stronger—standards used for human-written code: understand the change, verify it against the specification, run tests and automated checks, review security and maintainability, and keep a human accountable for the final decision.
[1] Anthropic — How Claude Remembers Your Project
[2] GitHub Docs — Best Practices for Using GitHub Copilot
[3] Anthropic — Best Practices for Claude Code
[4] Model Context Protocol — What Is MCP?
[5] Model Context Protocol — Connect to Remote MCP Servers
[6] Model Context Protocol — Security Best Practices
[7] GitHub Docs — Application Card: GitHub Copilot AgentsH