Adopt

AI Agents for Software Development: Delegating Coding and Tests

・ Employee Store Operations

Summary

Coding AI agents can read code, rewrite it, run tests and propose changes. But if you do not decide where they run and what they can touch, you risk leaking secrets and lowering quality. This article explains how to choose the tasks you delegate and the rules your team should set.

The tool features in this article were checked in each company's official documentation on October 2, 2026. Names, features and pricing models may change. Before you rely on them, check the latest information through the sources at the end.

When you use AI agents in software development, do not hand over all of development. Start with tasks that have a clear scope. A person always reviews the changes, and only changes that pass the tests get merged. With that as the baseline, let's look at which tasks to delegate and how the tools differ.

Tasks That Are Easy to Delegate: Tests, Fixes and Research

The easiest tasks to delegate are ones where a machine can confirm whether they are done. Writing tests, fixing the cause of a failing test, and conforming code to a set style can all be checked by whether tests and linters pass. Test automation with AI agents is an easy place to start, whether in contract system development or in-house development.

  • Adding tests: add tests to existing code and reduce the parts that are not covered
  • Fixing small bugs: fix issues that include steps to reproduce and the expected behavior
  • Research: find out where and what the codebase does, and write explanations or implementation plans
  • Updating documentation: revise the docs to match code changes

GitHub's official documentation also lists what Copilot's cloud agent can do: research a repository, create implementation plans, fix bugs, add small features, improve test coverage, update documentation, address technical debt, and resolve merge conflicts.

Harder to delegate are tasks without a settled specification and tasks where mistakes have a large impact. For authentication, payments, permissions, and processes that change production data, people do the design and review thoroughly, even if the AI writes a draft.

How Coding AI Agents Work

A coding AI agent works by repeating three actions: reading files, rewriting files, and running commands. It runs the tests, reads the results, and fixes the code if they fail, all on its own.

Where they run falls into two broad groups: on the developer's computer and in a cloud environment. Where an agent runs changes which files it can touch and when people check its work.

AI agents that run locally vs. in the cloud

Local (editor or terminal)

  • Rewrites files on the developer's computer directly
  • The developer watches and steers in real time
  • Scope is limited through settings and approvals

Cloud (GitHub and similar environments)

  • Works independently in a prepared environment
  • Results arrive as branches or pull requests
  • People review in the pull request

How the Main Tools Differ, Based on Official Information

Here we compare three tools, GitHub Copilot, OpenAI's Codex and Cursor, based only on what their official documentation says. The focus is not which is better, but where each one runs and how it is built for people to check its work.

How coding AI agents differ (checked October 2026)
Point of comparisonGitHub CopilotCodexCursor
Where it runsThe cloud agent runs in a GitHub Actions-powered environment. Agent mode in the IDE runs locallyCLI, IDE extension, cloudAgent inside the editor. There is also an Agent that runs in the cloud
Default safety settingsThe cloud agent works with one branch and one pull request per taskBy default, the CLI and IDE extension have no network access, and writes are limited to the working folderTerminal commands require approval by default
How to pass project conventionsCustom instruction files stored in the repositoryAGENTS.mdRules (.cursor/rules) or AGENTS.md

GitHub Copilot's Cloud Agent

Copilot's cloud agent is available on all paid Copilot plans. On Copilot Business and Copilot Enterprise, an administrator must enable the policy. While it works, it uses a temporary development environment powered by GitHub Actions, where it can explore code, make changes, and run automated tests and linters. A single session can run for up to 59 minutes. Costs draw on GitHub Actions minutes and AI credits.

Codex

Codex is available as a CLI in the terminal, as an editor extension, and in the cloud. The IDE extension supports VS Code and compatible editors such as Cursor and Windsurf. The CLI and IDE extension use OS-level sandboxing; by default they have no network access, and writes are limited to the working folder. They ask for approval before editing outside the working folder or running commands that need the network.

Cursor

Cursor's Agent can search the codebase, edit files and run terminal commands. Reading and searching files needs no approval. Files in the working folder, except configuration files, can be rewritten without approval and are saved to disk immediately. The official documentation recommends using version control so you can roll changes back.

If you want to use an AI agent in VS Code, you can choose Copilot's agent mode, Codex's IDE extension, and others. Anthropic's Claude Code is also a coding agent that works in the terminal, IDEs, a desktop app and the browser. We explain it in Building an AI Employee with Claude Code. Choose based on the editors your team uses, where your code lives, and whether the approval mechanism fits your company's rules.

Rules for Handling Code and Secrets

AI agents send the code you give them, and the contents of files they can read, to an AI model. Your team should decide three things: what the agent must not read, what must not be used for training, and what it must not run without approval.

  • Decide which files it must not read: files containing API keys or passwords, and customer data. In Cursor, .cursorignore excludes files from what the agent reads
  • Check the settings that keep code out of training: with Cursor's Privacy Mode on, code is not used for training by Cursor or the model providers. For teams, Privacy Mode is on by default
  • Limit network access and execution: decide as a team before loosening Codex's sandbox or Cursor's approval settings beyond the defaults
  • Approve connections to external tools: in Cursor, MCP connections need approval, and even after connecting, each tool call asks for approval

Secrets should not be written in code; keep them in environment variables or a secrets management system. During review, also check that code written by the AI agent does not contain keys written directly. We explain how to think about AI agent permissions in Managing AI Agent Permissions, and overall security in AI Agent Security.

Protect Quality with Review and Tests

Review changes made by an AI agent through the same process as changes written by people. Do not skip review or tests because the AI made the change. Asking for small, separate changes makes them easier to review.

From request to merging an AI agent's change
  1. 1Write the issueSteps to reproduce and completion criteria
  2. 2The AI does the workChanges code on a branch and runs tests
  3. 3Automated checksTests, linters and security scans
  4. 4A person reviewsCheck design, permissions and secrets
  5. 5MergeA person decides whether to merge

Write the completion criteria in the issue in a form that tests can confirm. For example, state which input should produce which result. If the criteria are vague, the AI may make changes that only exist to pass the tests.

When reviewing web application code, you can use 'How to Secure Your Website' from IPA, Japan's Information-technology Promotion Agency. The revised 7th edition covers 11 types of vulnerabilities, including SQL injection, OS command injection and cross-site scripting, and includes a security implementation checklist at the end. You can check AI-written code against this checklist too. When your organization sets rules for AI use, the 'AI Guidelines for Business' from Japan's Ministry of Internal Affairs and Communications (MIC) and Ministry of Economy, Trade and Industry (METI) are also a useful reference. Version 1.2 was published on March 31, 2026, and it also includes a checklist.

Employee Store's Development and Data Category

Employee Store is a marketplace where companies can adopt AI agents built by developers. One of its categories is development and data. Sellers can list in any format, including n8n or Dify workflows, their own custom-built agents, and agents that use the Claude or OpenAI APIs.

Pricing is one-time purchase, monthly, or setup fee plus monthly, and payment is through Stripe. After purchase, you message the seller in the deal room, check the deliverables, and accept them. Delivery is, in principle, within 3 business days, and you are refunded if it runs late. Before you adopt an agent, check on the listing page which repositories and tools it needs permission to access.

FAQ

If an AI agent writes the code, do I still need review?
Yes. Review changes made by an AI agent through the same process as changes written by people, and merge only changes that pass the tests. Changes that involve authentication, payments, permissions or production data need especially thorough human checks.
Are there AI agents I can use in VS Code?
GitHub Copilot's agent mode and Codex's IDE extension work in VS Code. Claude Code also works in IDEs. Choose based on your team's editors, where your code lives, and whether the approval mechanism fits your company's rules.
I am worried our internal code will be used to train AI.
Settings and conditions differ by tool. For example, in Cursor, code is not used for training when Privacy Mode is on. Check the official documentation of the tool you use for how it handles data and which settings your organization can enforce.

About the author

Employee Store OperationsThe operations team behind Employee Store, a marketplace for AI agents. We check tool features and pricing against official sources and list them at the end of each article. If you spot an error, please let us know via the contact form.

Sources

Ask AI

Ask AI if it fits your work.

Use your usual AI to explore what Employee Store offers and what to check before buying.

Opens an external AI service. Confirm pricing and deliverables on the listing page.